The End-to-End ML Workflow
A machine learning model is not a finished product — it is one component in a larger ML pipeline that must be built, versioned, deployed, and maintained. Most ML projects spend only a small fraction of their time on model training; the bulk of the engineering effort goes into data ingestion, feature computation, serving infrastructure, and monitoring.
A mature ML pipeline makes every step reproducible (the same inputs always produce the same outputs), observable (you can see what is happening in production), and maintainable (you can update individual stages without breaking others). These properties are not luxuries — they are prerequisites for deploying ML that actually stays reliable.
Data Versioning
Code version control (Git) is well understood, but datasets present a unique challenge: they can be gigabytes or terabytes, change frequently, and are often stored outside the repository. Without data versioning, reproducing a past experiment requires reconstructing the exact dataset that was used — often impossible in practice.
DVC (Data Version Control)
DVC extends Git to handle large files and datasets. Instead of storing data in the Git repository, DVC stores a small pointer file (a content hash) in Git and pushes the actual data to a remote storage backend (S3, GCS, Azure Blob, a local NAS, etc.). This gives you:
- Reproducibility — every Git commit can be linked to the exact dataset version used.
- Collaboration — teammates pull the data they need without checking large files into Git.
- Pipeline tracking — DVC can also version pipeline stages (preprocessing, training) and detect which stages need to re-run when inputs change.
The most common source of production ML failures: the feature computation during training differs from the feature computation at inference time. Lock your feature logic into versioned pipeline code that runs identically in both environments. Never compute features ad hoc at serving time.
Experiment Tracking
Every training run makes choices — which model architecture, which hyperparameters, which data version, which preprocessing steps. Without systematic logging, it becomes impossible to know which run produced which model or how to reproduce a past result. Experiment tracking captures every run's configuration, metrics, and outputs in a searchable database.
MLflow
MLflow is an open-source platform for the ML lifecycle. Its core components:
Weights & Biases (W&B)
Weights & Biases is a cloud-hosted experiment tracking platform particularly popular in deep learning. It provides rich visualization of training curves, hyperparameter sweeps via its Sweeps feature (which implements Bayesian optimization and random search), dataset versioning via Artifacts, and model registry. Its dashboard makes comparing dozens of runs across teams straightforward.
Model Serialization
A trained model must be saved in a format that can be loaded later for inference — possibly on a different machine, in a different language, or years later. The choice of serialization format affects portability, security, and inference speed.
Serialization Formats
Deployment
Serving a model means making its predictions available to other systems. The deployment strategy depends on latency requirements, request volume, and operational complexity tolerance.
Online vs. Batch Serving
Online serving exposes the model as a real-time API endpoint. Each request arrives, is preprocessed, passed to the model, and the prediction is returned — all within milliseconds. This is the right approach when predictions must be fresh (recommendation systems, fraud detection, real-time pricing).
Batch serving runs the model on a large dataset at a scheduled interval and stores predictions in a database that downstream systems can read. This is appropriate when predictions can be precomputed (daily email recommendations, offline scoring), and it is far simpler to operate than a real-time API.
Containerization and Orchestration
Packaging a model for deployment in a Docker container ensures that the Python version, library versions, and system dependencies are locked and reproducible across environments. Containers run identically on a laptop, a CI server, and a production Kubernetes cluster.
For high-traffic serving, models are deployed behind a load balancer with multiple replicas and horizontal autoscaling. Platforms like BentoML, Seldon Core, and KServe provide ML-specific serving infrastructure on top of Kubernetes — handling model loading, preprocessing, batching, and health checks.
Before routing real traffic to a new model, run it in "shadow mode": send the same requests to both the old and new model, log both predictions, but only return the old model's output to users. This lets you compare predictions and catch regressions in production data distributions before the new model has any user impact.
Monitoring and Retraining
Models degrade over time. The world changes — user behavior shifts, product features change, economic conditions evolve — and the patterns the model learned from historical data become less predictive of the current distribution. This is called concept drift (the relationship between features and labels changes) or data drift (the feature distribution changes). Both reduce model performance without any visible error in the system.
What to Monitor
- Prediction distribution — are the model's output scores or class probabilities shifting over time? A sudden change often signals a data quality issue or drift.
- Feature distributions — are the input features staying within the expected range and distribution? Use statistical tests (Population Stability Index, KS test) to detect drift.
- Business metrics — click-through rate, conversion rate, churn rate. These are the ground truth of whether the model is actually helping.
- Ground truth labels — when delayed labels become available (e.g., whether a flagged transaction was actually fraud), compute model accuracy on recent data.
Retraining Strategies
A/B Testing for Models
When a new model candidate is ready, how do you decide whether to promote it? The gold standard is an A/B test: route a fraction of live traffic to the new model (treatment group) and keep the rest on the old model (control group). Measure the business metric that matters — revenue, engagement, accuracy with real labels — and use a statistical significance test to determine whether the difference is real or noise.
Key discipline: define your success metric and minimum detectable effect before running the test, and commit to the planned sample size before looking at results. Peeking at partial results and stopping early inflates false-positive rates dramatically.
An ML pipeline connects data to production through reproducible, versioned stages. Use DVC for data versioning and MLflow or W&B for experiment tracking — never rely on manual notes. Serialize models in portable formats (ONNX, TorchScript, MLflow Model) and containerize serving environments. Monitor prediction and feature distributions for drift; trigger retraining when performance degrades. Validate new models with A/B tests on real traffic before full rollout.