Learning path

Full curriculum

Full curriculum

Unit content

Features, targets, and the design matrix

A machine-learning dataset usually separates information used to make a prediction from the quantity to be predicted.

For example, if each house is described by floor area, age and number of rooms, those quantities are features. If the goal is to predict sale price, price is the target.

For $n$ examples with $d$ numerical features, it is convenient to arrange the features in a design matrix

$$X=\begin{bmatrix}x_{11}&\cdots&x_{1d}\\vdots&&\vdots\x_{n1}&\cdots&x_{nd}\end{bmatrix},$$

where row $i$ describes example $i$ and column $j$ is feature $j$. The targets form a vector $y$ in ordinary regression or classification problems.

This notation separates three objects that should not be confused:

  • an example is one row;
  • a feature is one measured or constructed input variable;
  • a target is what supervised learning tries to predict.

A feature need not be a raw measurement. From a timestamp, for instance, one might construct hour-of-day or day-of-week features. Such transformations change the representation presented to the model without changing the underlying examples.

Thinking explicitly in terms of examples, features and targets makes later questions about preprocessing, leakage, dimensionality and model coefficients much clearer.