Learning path

Full curriculum

Full curriculum

Unit content

Categorical features and one-hot encoding

A categorical variable represents membership in a finite set of categories rather than a numerical quantity. Assigning arbitrary numbers to categories can accidentally introduce an ordering that the model treats as meaningful.

For a feature color with categories red, green and blue, one-hot encoding replaces it with indicator features such as

$$(1,0,0),\quad(0,1,0),\quad(0,0,1).$$

Each coordinate answers whether the example belongs to one category. The representation does not imply that blue is numerically larger than green or that red is twice anything.

For linear models with an intercept, one indicator is often omitted to avoid exact linear dependence among columns: if red, green and blue indicators always sum to one, all three plus an intercept are redundant.

One-hot encoding can become inefficient for variables with thousands or millions of categories. In those cases, hashing, learned embeddings or domain-specific representations may be preferable.

A fitted encoder includes its category vocabulary. When it is applied to additional examples, that same vocabulary must be reused, together with an explicit policy for previously unseen categories.