Learning path

Full curriculum

Full curriculum

Unit content

Least squares as a Gaussian noise model

Least-squares curve fitting can be derived from a probabilistic assumption rather than introduced only as a geometric rule.

Suppose observations follow

$$y_i=f(x_i;\theta)+\varepsilon_i,$$

where the errors are independent Gaussian variables with mean zero and common standard deviation $\sigma$:

$$\varepsilon_i\sim\mathcal N(0,\sigma^2).$$

The likelihood

For fixed $\sigma$, the likelihood of the observed outputs is proportional to

$$\exp\left(-\frac{1}{2\sigma^2}\sum_i(y_i-f(x_i;\theta))^2\right).$$

Maximizing this likelihood is therefore equivalent to minimizing

$$\sum_i(y_i-f(x_i;\theta))^2.$$

The familiar square in least squares is thus tied to an assumption of Gaussian observation noise.

Different noise, different loss

If the error distribution changes, maximum likelihood can lead to a different fitting objective. Absolute-error losses, for example, arise naturally from other noise models.

Models should explain residuals

A fitted curve is only part of the statistical model. Residual patterns can reveal that the assumed noise is not independent, has changing variance or that the mean function itself is misspecified.