Home / ML 101 / Module 10 / Lesson 3

ML Pipeline

From raw data to production — building reproducible, monitored, and maintainable machine learning systems.

~15 min read M10 · L3 Intermediate

The End-to-End ML Workflow

A machine learning model is not a finished product — it is one component in a larger ML pipeline that must be built, versioned, deployed, and maintained. Most ML projects spend only a small fraction of their time on model training; the bulk of the engineering effort goes into data ingestion, feature computation, serving infrastructure, and monitoring.

A mature ML pipeline makes every step reproducible (the same inputs always produce the same outputs), observable (you can see what is happening in production), and maintainable (you can update individual stages without breaking others). These properties are not luxuries — they are prerequisites for deploying ML that actually stays reliable.

01
Data Ingestion & Versioning
Collect raw data from sources; version datasets so experiments are reproducible.
02
Feature Engineering
Transform raw data into model-ready features; store feature logic to avoid train/serve skew.
03
Training & Experiment Tracking
Train models; log hyperparameters, metrics, and artifacts for every run.
04
Model Serialization & Registry
Save the trained model in a portable format; register it with metadata for deployment gating.
05
Serving & Deployment
Expose the model as an API or batch job; package dependencies for consistent inference.
06
Monitoring & Retraining
Track prediction quality and data drift in production; trigger retraining when performance degrades.

Data Versioning

Code version control (Git) is well understood, but datasets present a unique challenge: they can be gigabytes or terabytes, change frequently, and are often stored outside the repository. Without data versioning, reproducing a past experiment requires reconstructing the exact dataset that was used — often impossible in practice.

DVC (Data Version Control)

DVC extends Git to handle large files and datasets. Instead of storing data in the Git repository, DVC stores a small pointer file (a content hash) in Git and pushes the actual data to a remote storage backend (S3, GCS, Azure Blob, a local NAS, etc.). This gives you:

Train/Serve Skew — The Silent Killer

The most common source of production ML failures: the feature computation during training differs from the feature computation at inference time. Lock your feature logic into versioned pipeline code that runs identically in both environments. Never compute features ad hoc at serving time.

Experiment Tracking

Every training run makes choices — which model architecture, which hyperparameters, which data version, which preprocessing steps. Without systematic logging, it becomes impossible to know which run produced which model or how to reproduce a past result. Experiment tracking captures every run's configuration, metrics, and outputs in a searchable database.

MLflow

MLflow is an open-source platform for the ML lifecycle. Its core components:

MLflow Tracking
Experiment logging
Log parameters, metrics, and artifacts (models, plots) for every run via a simple API.
MLflow Projects
Reproducible runs
Package ML code with its dependencies so anyone can re-run experiments in any environment.
MLflow Models
Portable format
Standard model format with a flavor system supporting scikit-learn, PyTorch, TensorFlow, and more.
Model Registry
Deployment gating
Register models with lifecycle stages (Staging → Production → Archived) and metadata.

Weights & Biases (W&B)

Weights & Biases is a cloud-hosted experiment tracking platform particularly popular in deep learning. It provides rich visualization of training curves, hyperparameter sweeps via its Sweeps feature (which implements Bayesian optimization and random search), dataset versioning via Artifacts, and model registry. Its dashboard makes comparing dozens of runs across teams straightforward.

What to Log Every Run
\text{Run} = \{\text{git\_sha},\, \text{data\_hash},\, \theta,\, \mathcal{L}_{\text{val}}\}
A complete run record links code version, data version, hyperparameters, and metrics — forming the provenance chain needed to reproduce any result.

Model Serialization

A trained model must be saved in a format that can be loaded later for inference — possibly on a different machine, in a different language, or years later. The choice of serialization format affects portability, security, and inference speed.

Serialization Formats

joblib
scikit-learn models
Efficient Python serialization for numpy arrays and sklearn pipelines. Convenient for trusted internal use; never load files from untrusted sources.
ONNX
Cross-platform
Open Neural Network Exchange: export from PyTorch/sklearn, run in any ONNX Runtime (C++, Java, .NET, JavaScript).
TorchScript / SavedModel
Deep learning
PyTorch's TorchScript and TensorFlow's SavedModel compile the model for optimized serving without a Python interpreter.
MLflow Model
Multi-framework
Wraps any model with a standard interface. Supports multiple "flavors" so the same model can be loaded via different backends.

Deployment

Serving a model means making its predictions available to other systems. The deployment strategy depends on latency requirements, request volume, and operational complexity tolerance.

Online vs. Batch Serving

Online serving exposes the model as a real-time API endpoint. Each request arrives, is preprocessed, passed to the model, and the prediction is returned — all within milliseconds. This is the right approach when predictions must be fresh (recommendation systems, fraud detection, real-time pricing).

Batch serving runs the model on a large dataset at a scheduled interval and stores predictions in a database that downstream systems can read. This is appropriate when predictions can be precomputed (daily email recommendations, offline scoring), and it is far simpler to operate than a real-time API.

Containerization and Orchestration

Packaging a model for deployment in a Docker container ensures that the Python version, library versions, and system dependencies are locked and reproducible across environments. Containers run identically on a laptop, a CI server, and a production Kubernetes cluster.

For high-traffic serving, models are deployed behind a load balancer with multiple replicas and horizontal autoscaling. Platforms like BentoML, Seldon Core, and KServe provide ML-specific serving infrastructure on top of Kubernetes — handling model loading, preprocessing, batching, and health checks.

Shadow Mode Deployment

Before routing real traffic to a new model, run it in "shadow mode": send the same requests to both the old and new model, log both predictions, but only return the old model's output to users. This lets you compare predictions and catch regressions in production data distributions before the new model has any user impact.

Monitoring and Retraining

Models degrade over time. The world changes — user behavior shifts, product features change, economic conditions evolve — and the patterns the model learned from historical data become less predictive of the current distribution. This is called concept drift (the relationship between features and labels changes) or data drift (the feature distribution changes). Both reduce model performance without any visible error in the system.

What to Monitor

Retraining Strategies

Scheduled Retraining
Periodic
Retrain on a fixed schedule (daily, weekly). Simple to implement but may retrain unnecessarily or not quickly enough after drift.
Triggered Retraining
Drift-based
Automatically trigger a retraining run when monitoring detects significant drift or performance degradation.
Continuous Training
Online learning
Update model weights incrementally as new data arrives. Requires careful stability controls to prevent catastrophic forgetting.
Human-in-the-Loop
Manual approval
Retraining is automated but a human reviews and approves the new model before it replaces the current production model.

A/B Testing for Models

When a new model candidate is ready, how do you decide whether to promote it? The gold standard is an A/B test: route a fraction of live traffic to the new model (treatment group) and keep the rest on the old model (control group). Measure the business metric that matters — revenue, engagement, accuracy with real labels — and use a statistical significance test to determine whether the difference is real or noise.

Key discipline: define your success metric and minimum detectable effect before running the test, and commit to the planned sample size before looking at results. Peeking at partial results and stopping early inflates false-positive rates dramatically.

A/B Test Sample Size
n = \frac{2\sigma^2(z_{\alpha/2} + z_\beta)^2}{\delta^2}
n: required users per group. σ²: outcome variance. δ: minimum effect size worth detecting. z values for α=0.05, power=0.8: z_{α/2}=1.96, z_β=0.84.

Key Takeaways

An ML pipeline connects data to production through reproducible, versioned stages. Use DVC for data versioning and MLflow or W&B for experiment tracking — never rely on manual notes. Serialize models in portable formats (ONNX, TorchScript, MLflow Model) and containerize serving environments. Monitor prediction and feature distributions for drift; trigger retraining when performance degrades. Validate new models with A/B tests on real traffic before full rollout.

Previous M10-L2: Model Selection & Tuning Module Overview Next Lesson M11-L1: Training at Scale