ML 101
M10 · L03
Practical ML

The ML Pipeline

From raw data to production — building reproducible, monitored, and maintainable machine learning systems that stay reliable over time.

01 / 13
ML 101
M10 · L03
End-to-End Workflow

Six Stages

  • Data ingestion & versioning
  • Feature engineering
  • Training & experiment tracking
  • Model serialization & registry
  • Serving & deployment
  • Monitoring & retraining
02 / 13
ML 101
M10 · L03
Data Versioning

DVC — Git for Data

DVC stores a content hash pointer in Git and pushes actual data to remote storage (S3, GCS, Azure Blob). Every Git commit links to the exact dataset version used.

Key benefit
Teammates pull only the data they need. Pipeline stages detect when inputs change and re-run only what is necessary.
03 / 13
ML 101
M10 · L03
The Silent Killer

Train/Serve Skew

The most common production ML failure: feature computation differs between training and inference. Lock feature logic into versioned pipeline code that runs identically in both environments.

Rule
Never compute features ad hoc at serving time. One code path, one version, two environments.
04 / 13
ML 101
M10 · L03
Experiment Tracking

MLflow & W&B

  • MLflow Tracking — log params, metrics, and artifacts per run
  • MLflow Registry — lifecycle stages: Staging → Production → Archived
  • W&B Sweeps — Bayesian hyperparameter search across runs
  • W&B Artifacts — dataset and model versioning in the cloud
  • Every run: git sha + data hash + hyperparameters + validation metric
05 / 13
ML 101
M10 · L03
Run Provenance

A Complete Record

Every training run must record the full provenance chain so any result can be reproduced:

Run Record
\text{Run}=\{\text{git\_sha},\,\text{data\_hash},\,\theta,\,\mathcal{L}_{\text{val}}\}
06 / 13
ML 101
M10 · L03
Model Serialization

Portable Formats

  • joblib — efficient for sklearn pipelines; trusted internal use only
  • ONNX — cross-platform: run in C++, Java, .NET, JavaScript
  • TorchScript — compiled PyTorch; no Python interpreter needed at inference
  • MLflow Model — wraps any framework; standard flavors interface
07 / 13
ML 101
M10 · L03
Deployment

Online vs. Batch

Online
Real-time API · ms latency
Batch
Scheduled · precomputed

Package as a Docker container for reproducible environments. Use BentoML, Seldon Core, or KServe for ML-specific serving on Kubernetes.

08 / 13
ML 101
M10 · L03
Safe Rollout

Shadow Mode

Send live traffic to both old and new models simultaneously. Return the old model's output to users, but log both predictions for comparison. Catch regressions before the new model has any user impact.

After shadow mode
Graduate to A/B test → canary deployment → full rollout. Each stage increases exposure gradually.
09 / 13
ML 101
M10 · L03
Monitoring

Drift & Degradation

  • Prediction distribution — are output scores shifting over time?
  • Feature distributions — PSI or KS test to detect data drift
  • Business metrics — conversion rate, engagement, churn: the real ground truth
  • Delayed labels — compute actual accuracy once ground truth arrives
10 / 13
ML 101
M10 · L03
A/B Testing

Validate with Real Traffic

Split live traffic between old (control) and new (treatment) model. Required sample size per group:

Sample Size
n=\frac{2\sigma^2(z_{\alpha/2}+z_\beta)^2}{\delta^2}

Define your metric and minimum detectable effect before running. Never peek early.

11 / 13
ML 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

12 / 13
ML 101
Key Takeaways
Summary

Key Takeaways

  • ML pipelines need reproducibility, observability, and maintainability
  • DVC versions datasets; MLflow / W&B track experiments
  • Avoid train/serve skew: one feature code path for both environments
  • Serialize in portable formats (ONNX, TorchScript); containerize serving
  • Monitor prediction and feature distributions for drift; retrain proactively
  • Validate new models with A/B tests before full rollout
13 / 13