Learning path

Full curriculum

Full curriculum

Unit content

Data leakage in machine learning

Data leakage occurs when model training or model selection uses information that would not legitimately be available when making predictions on new cases.

A subtle example is preprocessing. Suppose every numerical feature is standardized using the mean and standard deviation computed from the entire dataset, and only afterward the data are split into training and test sets. The test examples have influenced the transformation applied to the training data, so the reported test performance is no longer fully out of sample.

The safe pattern is:

  1. split the data by its intended evaluation structure;
  2. fit preprocessing operations using only training data;
  3. apply the fitted transformations to validation and test data without refitting them.

Leakage can also be semantic. Predicting whether a patient will be readmitted using a variable recorded only after discharge gives the model information from the future. Randomly splitting repeated measurements from the same person can leak identity-specific patterns between train and test sets.

The central test is: could this information have been known at the moment the prediction is supposed to be made? If not, using it makes evaluation unrealistically optimistic.