Unit content
Linear least-squares optimization
When a linear system
$$Ax=b$$
has no exact solution, least squares chooses $x$ to minimize the squared residual:
$$\min_x \lVert Ax-b\rVert^2.$$
At an optimum, the residual
$$r=b-Ax$$
is orthogonal to every column of $A$. Therefore
$$A^T(b-Ax)=0,$$
which gives the normal equations
$$A^TAx=A^Tb.$$
Geometrically, $Ax$ is the orthogonal projection of $b$ onto the column space of $A$.
If the columns of $A$ are linearly independent, the least-squares minimizer is unique.
The normal equations explain the mathematics, but numerical software often uses factorizations that avoid explicitly forming $A^TA$, because that transformation can amplify sensitivity to finite-precision errors.