Unit content
Maximum-margin classifiers and support vector machines
A linear classifier can separate two classes with many different decision boundaries. A maximum-margin classifier prefers the boundary that leaves the widest geometric gap between the classes.
For labels $y_i\in{-1,+1}$ and linear score $w^Tx+b$, the hard-margin support vector machine solves
$$\min_{w,b}\frac12|w|^2$$
subject to
$$y_i(w^Tx_i+b)\ge 1$$
for every training example. Minimizing $|w|$ maximizes the margin after this normalization.
Only the examples touching the margin constraints determine the optimum. These are the support vectors.
Real data are rarely perfectly separable. A soft-margin SVM introduces slack through the hinge loss
$$\max(0,1-y(w^Tx+b))$$
and balances margin width against violations using a regularization parameter.
The geometric viewpoint explains why feature scaling matters: rescaling one coordinate changes distances and therefore changes the margin geometry. SVMs can also use kernels to create nonlinear decision boundaries without explicitly constructing every transformed feature.