Learning path

Full curriculum

Full curriculum

Unit content

An end-to-end supervised machine-learning workflow

A reliable supervised-learning project is a sequence of information boundaries, not merely a call to a fitting algorithm.

Suppose the goal is to predict customer churn from account data.

  1. Define the prediction point and target. Features must be information available before the decision; later information would leak the future.
  2. Reserve evaluation data. Split by the deployment structure—for example by customer and, when appropriate, by time.
  3. Build preprocessing from training data. Impute missing values, encode categories and scale features as required by the candidate models.
  4. Establish baselines and metrics. For rare churn, accuracy alone is inadequate; precision-recall behavior may matter more.
  5. Fit justified candidate model families. Different hypothesis spaces encode different assumptions about the problem and should be compared under the same evaluation protocol.
  6. Choose hyperparameters using validation or cross-validation. Every fitted preprocessing step belongs inside this selection loop.
  7. Inspect errors, probability quality and relevant subgroup behavior. Revise features or model assumptions using development data only.
  8. Evaluate the frozen procedure once on the test set. This estimates performance after all choices have been made.
  9. Deploy and monitor distribution shift. Future data can stop resembling the evaluation setting.

The important object being evaluated is the entire procedure—representation, preprocessing, fitting, selection and decision rule—not just the mathematical model at its center.