Learning path

Full curriculum

Full curriculum

Unit content

Regularized optimization and penalty terms

An optimization objective can include a penalty that expresses preference among otherwise feasible solutions:

$$\min_\theta; L(\theta)+\lambda R(\theta),$$

where $L$ measures the original objective, $R$ is the penalty and $\lambda\ge0$ controls the trade-off.

An $L_2$ penalty

$$R(\theta)=\lVert\theta\rVert_2^2$$

discourages large parameter magnitudes. An $L_1$ penalty

$$R(\theta)=\lVert\theta\rVert_1$$

can favor sparse solutions with many zero components.

Changing $\lambda$ changes the optimization problem: stronger regularization accepts more data-fit error in exchange for satisfying the preference encoded by the penalty.

Regularization can improve robustness, resolve underdetermined problems or encode prior structural preferences. Its meaning comes from what $R$ rewards or discourages, not merely from adding another term to the formula.