Unit content
Missing data and imputation for machine learning
Real datasets often contain missing feature values. A learning algorithm usually cannot treat an absent number as if it were an ordinary numeric value, so the missingness must be represented deliberately.
One approach is imputation: replace a missing value with an estimate such as a reference-set median. For example, if observed ages in the reference data have median 37, missing ages can be replaced by 37.
Imputation does not recreate the unknown value. It only supplies a usable representation. It is often helpful to add a binary indicator such as age_missing, because the fact that a value is absent may itself carry information.
Different missingness mechanisms matter. A laboratory measurement may be absent because a test was not ordered, while a sensor may fail randomly. A simple imputation rule can behave very differently in those cases.
A fitted imputer includes any statistics it estimated, such as medians, and those values should be reused when transforming additional examples. The same missing-value representation must be available whenever the learned model is later applied.