The ML Pipeline
From raw data to production — building reproducible, monitored, and maintainable machine learning systems that stay reliable over time.
Six Stages
- Data ingestion & versioning
- Feature engineering
- Training & experiment tracking
- Model serialization & registry
- Serving & deployment
- Monitoring & retraining
DVC — Git for Data
DVC stores a content hash pointer in Git and pushes actual data to remote storage (S3, GCS, Azure Blob). Every Git commit links to the exact dataset version used.
Train/Serve Skew
The most common production ML failure: feature computation differs between training and inference. Lock feature logic into versioned pipeline code that runs identically in both environments.
MLflow & W&B
- MLflow Tracking — log params, metrics, and artifacts per run
- MLflow Registry — lifecycle stages: Staging → Production → Archived
- W&B Sweeps — Bayesian hyperparameter search across runs
- W&B Artifacts — dataset and model versioning in the cloud
- Every run: git sha + data hash + hyperparameters + validation metric
A Complete Record
Every training run must record the full provenance chain so any result can be reproduced:
Portable Formats
- joblib — efficient for sklearn pipelines; trusted internal use only
- ONNX — cross-platform: run in C++, Java, .NET, JavaScript
- TorchScript — compiled PyTorch; no Python interpreter needed at inference
- MLflow Model — wraps any framework; standard flavors interface
Online vs. Batch
Package as a Docker container for reproducible environments. Use BentoML, Seldon Core, or KServe for ML-specific serving on Kubernetes.
Shadow Mode
Send live traffic to both old and new models simultaneously. Return the old model's output to users, but log both predictions for comparison. Catch regressions before the new model has any user impact.
Drift & Degradation
- Prediction distribution — are output scores shifting over time?
- Feature distributions — PSI or KS test to detect data drift
- Business metrics — conversion rate, engagement, churn: the real ground truth
- Delayed labels — compute actual accuracy once ground truth arrives
Validate with Real Traffic
Split live traffic between old (control) and new (treatment) model. Required sample size per group:
Define your metric and minimum detectable effect before running. Never peek early.
Key Takeaways
- ML pipelines need reproducibility, observability, and maintainability
- DVC versions datasets; MLflow / W&B track experiments
- Avoid train/serve skew: one feature code path for both environments
- Serialize in portable formats (ONNX, TorchScript); containerize serving
- Monitor prediction and feature distributions for drift; retrain proactively
- Validate new models with A/B tests before full rollout