Learning path

Full curriculum

Full curriculum

Unit content

Ridge and lasso regression

Linear models can overfit when many coefficients are weakly constrained by the data. Ridge and lasso regression control this by adding penalties to the least-squares objective.

Ridge regression minimizes

$$|y-X\beta|_2^2+\lambda|\beta|_2^2,$$

where $\lambda\ge 0$ controls the strength of the penalty. It tends to shrink coefficients smoothly toward zero and is especially useful when predictors are correlated.

Lasso regression instead uses an $L_1$ penalty:

$$|y-X\beta|_2^2+\lambda|\beta|_1.$$

Because the $L_1$ geometry has corners, lasso solutions often set some coefficients exactly to zero, producing sparse models.

For example, with hundreds of weakly useful predictors, ordinary least squares may fit unstable large coefficients. Increasing $\lambda$ trades training fit for a simpler coefficient vector that can generalize better.

The regularization strength is a hyperparameter and must be chosen from held-out data or cross-validation, not by selecting the value that minimizes training loss. Feature scaling is also important because the penalty acts directly on coefficient magnitudes.