Unit content
Feature engineering and basis expansion
A model can become much more expressive when raw inputs are transformed into features that expose useful structure.
Suppose a scalar input $x$ has a curved relationship with target $y$. A linear predictor using features
$$\phi(x)=(1,x,x^2)$$
fits
$$\hat y=\beta_0+\beta_1x+\beta_2x^2,$$
which is nonlinear in $x$ while remaining linear in the fitted coefficients.
Other feature engineering examples include interaction terms $x_1x_2$, logarithms of positive skewed variables, cyclic encodings of time, or physically meaningful ratios.
The transformation encodes assumptions about which distinctions should be easy for the model to express. This can improve learning dramatically when the representation matches the problem.
But feature construction also enlarges the hypothesis space. Adding many arbitrary polynomial and interaction features can increase variance and create strong correlations among predictors. Complexity control and held-out evaluation therefore become especially important.
Feature engineering and representation learning solve the same broad problem at different levels: choose coordinates in which the desired pattern is easier to model.