Learning path

Full curriculum

Full curriculum

Unit content

Numerical feature scaling and standardization

Many learning algorithms compare feature magnitudes directly, so the numerical scale of each feature can materially affect training.

A common transformation is standardization. For a feature with reference mean $\mu$ and standard deviation $\sigma$,

$$z=\frac{x-\mu}{\sigma}.$$

After transformation, the reference feature has mean approximately zero and standard deviation one.

Suppose one model uses age measured in years and annual income measured in euros. A Euclidean distance or a gradient step can be dominated by the income coordinate simply because its numerical values are much larger. Standardization puts both coordinates on comparable scales without claiming that they are equally informative.

Scaling is especially important for distance-based methods, margin methods and gradient-based optimization. Tree-based models, by contrast, usually depend on orderings and thresholds rather than absolute scale and therefore need it much less.

A fitted scaling transformation consists not only of the formula but of its reference values $\mu$ and $\sigma$. When the transformation is applied to additional examples, those same values must be reused if the resulting coordinates are to remain comparable.