Learning path

Full curriculum

Full curriculum

Unit content

Maximum-margin classifiers and support vector machines

A linear classifier can separate two classes with many different decision boundaries. A maximum-margin classifier prefers the boundary that leaves the widest geometric gap between the classes.

For labels $y_i\in{-1,+1}$ and linear score $w^Tx+b$, the hard-margin support vector machine solves

$$\min_{w,b}\frac12|w|^2$$

subject to

$$y_i(w^Tx_i+b)\ge 1$$

for every training example. Minimizing $|w|$ maximizes the margin after this normalization.

Only the examples touching the margin constraints determine the optimum. These are the support vectors.

Real data are rarely perfectly separable. A soft-margin SVM introduces slack through the hinge loss

$$\max(0,1-y(w^Tx+b))$$

and balances margin width against violations using a regularization parameter.

The geometric viewpoint explains why feature scaling matters: rescaling one coordinate changes distances and therefore changes the margin geometry. SVMs can also use kernels to create nonlinear decision boundaries without explicitly constructing every transformed feature.