From Data to Predictions
You now know what machine learning is and the three main types — supervised, unsupervised, and reinforcement learning. But how does a machine learning model actually learn? In this lesson, we'll build your first mental model of the ML process — from raw data to trained predictions.
A machine learning model is a function that maps inputs to outputs. Instead of a human writing the rules, the model discovers them from data through an iterative process of prediction, error measurement, and adjustment.
Traditional programming: rules + data → answers | ML: data + answers → rulesWhat is a Model?
In machine learning, a model is simply a mathematical function. You feed it inputs — like square footage and number of bedrooms — and it produces an output — like a predicted house price. The model's job is to learn the right mapping from examples.
Think of it as a black box with adjustable knobs inside. At first, the knobs are set randomly and the predictions are terrible. Through training, the model tunes these knobs (called parameters or weights) until its predictions become accurate.
Features & Labels
Every supervised ML problem has two ingredients. Features (often called X) are the input variables — the information the model uses to make predictions. Labels (often called y) are the answers we want the model to predict.
In a spam detector, the features might be word frequencies, sender reputation, and email length. The label is binary: spam or not spam. The quality of your features often matters more than the choice of algorithm. Garbage in, garbage out.
Training Data
A model learns from training data — a collection of examples where we already know both the features and the correct labels. The more examples, and the more representative they are of real-world scenarios, the better the model learns.
Think of it like studying for an exam: more practice problems covering more topics leads to better performance. But if your practice problems only cover half the syllabus, you'll fail the questions on the other half. Biased or incomplete training data leads to biased or incomplete models.
The Learning Loop
Learning happens through an iterative loop with three steps:
1. Predict: The model uses its current parameters to make predictions on the training data.
2. Measure error: A loss function compares the predictions to the actual answers and produces a single number measuring how wrong the model is.
3. Adjust: An optimization algorithm tweaks the parameters to reduce the error.
This cycle repeats thousands or millions of times. With each iteration, the model gets slightly better. This is the fundamental mechanism behind all supervised learning.
Loss Functions
The loss function (or cost function) measures how wrong the model's predictions are. It's the model's report card. For regression problems — predicting continuous values like price — the most common loss is Mean Squared Error:
Gradient Descent
Gradient descent is how the model adjusts its parameters. Imagine standing on a hilly landscape in fog — you can't see the bottom, but you can feel the slope under your feet. You take a step downhill, then another, and another. Eventually you reach a valley.
That's gradient descent: compute the slope (gradient) of the loss with respect to each parameter, then take a small step in the direction that reduces the loss. The size of each step is the learning rate — too large and you overshoot; too small and training takes forever.
Linear Regression: The Simplest Model
The simplest ML model is linear regression: a straight line that best fits the data.
Despite its simplicity, linear regression is widely used and powerful. More importantly, every concept we've covered — features, loss functions, gradient descent — applies in exactly the same way to neural networks with millions of parameters. Lesson 2.1 works through linear regression in full — the cost function, gradient descent, and a closed-form solution — where the bias b is written w₀.
Overfitting vs Underfitting
Overfitting is when a model memorizes the training data but fails on new examples. It has learned the noise, not the signal. Underfitting is when the model is too simple to capture the real patterns in the data.
The sweet spot is generalization — learning the true underlying pattern without memorizing specific examples. This is the central challenge of machine learning, and much of ML practice is about finding this balance.
Train/Test Split
To detect overfitting, we hold out some data that the model never sees during training — the test set. We train on the training set and evaluate on the test set. If the model performs well on training but poorly on test data, it's overfitting.
A typical split is 80% training, 20% testing. This simple technique is the foundation of all model evaluation in machine learning. Without it, you have no idea whether your model actually works on real data.
1. Collect & prepare data → 2. Split into train/test → 3. Train the model → 4. Evaluate on test set → 5. Iterate if needed. Machine learning is inherently iterative.
- A model is a function that learns input-output mappings from data. Features (X) are inputs; labels (y) are outputs.
- The learning loop: predict → measure error (loss function) → adjust parameters (gradient descent). Repeat.
- Mean Squared Error measures regression accuracy; gradient descent optimizes parameters by following the downhill slope of the loss.
- Linear regression (y = wx + b) is the simplest model but illustrates all core ML concepts.
- Overfitting memorizes noise; underfitting misses patterns. Train/test splits detect overfitting by evaluating on unseen data.