Unit content
Principal component analysis
Principal component analysis (PCA) finds orthogonal directions along which centered data vary most strongly, providing a lower-dimensional linear representation.
Let centered observations be rows of matrix $X$. The sample covariance matrix is proportional to
$$X^TX.$$
Its eigenvectors give the principal directions. The eigenvector with the largest eigenvalue is the first principal component direction; it maximizes the variance of the projected data among all unit directions.
If $v_1$ is that direction, each centered example $x$ receives one-dimensional coordinate
$$z_1=x^Tv_1.$$
Keeping the first $r$ eigenvectors $V_r$ gives the reduced representation
$$z=V_r^Tx.$$
The associated eigenvalues quantify how much variance each component explains.
For a two-dimensional cloud stretched along a diagonal, PCA rotates the coordinate system so the first axis follows that long direction. Keeping only that coordinate compresses the data while preserving much of its variation.
PCA is sensitive to feature scale and captures only linear structure. High variance is also not automatically the same as predictive or semantically important information.