Unit content
Unsupervised learning
In unsupervised learning, the data provide inputs but no externally supplied target label for each example. The goal is to discover or construct useful structure in the inputs themselves.
If the dataset is
$$x_1,\ldots,x_n,$$
an unsupervised method might:
- group similar examples into clusters;
- find a lower-dimensional representation;
- estimate the probability distribution that generated the data;
- detect unusual examples.
For instance, a retailer may have customer purchase vectors without predefined customer types. One method might propose groups with similar behavior, while another might compress many correlated measurements into a few informative coordinates.
Unlike supervised learning, there is often no single ground-truth target against which every result can be scored. Evaluation therefore depends more strongly on the purpose of the representation and on structural criteria.
Unsupervised learning does not mean learning without assumptions. Distance metrics, distribution families, number of groups and representation choices all encode inductive biases about which structures should count as meaningful.