Learning path

Full curriculum

Full curriculum

Unit content

The curse of dimensionality in machine learning

As dimension grows, geometric intuition from low-dimensional spaces can fail in ways that make learning harder. This collection of effects is called the curse of dimensionality.

Consider a unit cube in $d$ dimensions. A smaller cube whose side length is $r$ occupies fraction

$$r^d$$

of the total volume. To capture 10% of the volume, its side length must be

$$r=0.1^{1/d}.$$

In one dimension this is $0.1$, but in 100 dimensions it is about $0.977$. A neighborhood containing a modest fraction of the data must therefore extend across almost the entire range of every coordinate.

Consequences include:

  • local neighborhoods become data-hungry;
  • distances between points often become less discriminative;
  • flexible models have vastly more ways to fit noise;
  • density estimation requires much more data.

This is why local distance-based methods can degrade when many irrelevant dimensions are added and why dimensionality reduction, regularization and informative feature design can help.

High dimensionality is not inherently fatal—structured data can lie near much lower-dimensional manifolds—but dimension increases the amount of structure a learner must exploit rather than estimate freely.