The Three Paradigms
In the previous lesson, we learned that machine learning is about finding patterns in data. But not all learning is the same. Just as humans learn differently — from teachers, from exploration, and from trial and error — machine learning algorithms fall into distinct paradigms, each suited to different kinds of problems.
The three main paradigms are supervised learning, unsupervised learning, and reinforcement learning. Understanding when to use each one is one of the most important skills in applied ML. Let us explore each in depth.
Supervised = learning with a teacher who provides the answers. Unsupervised = exploring a library with no catalog and organizing the books yourself. Reinforcement = learning to ride a bike by falling down and getting back up.
Supervised Learning
Supervised learning is the most widely used paradigm in machine learning. The idea is simple: you provide the algorithm with labeled training data — pairs of inputs and their correct outputs — and the model learns a function that maps inputs to outputs. Once trained, it can predict the output for new, unseen inputs.
Supervised learning problems fall into two categories. Classification predicts a discrete label — is this email spam or not? Is this tumor malignant or benign? Regression predicts a continuous value — what will this house sell for? How much revenue will we generate next quarter?
The training process works by minimizing a loss function — a measure of how far the model's predictions are from the true labels. For regression, the most common loss function is Mean Squared Error:
Common supervised learning algorithms include linear regression for predicting continuous values, decision trees that split data based on feature thresholds, support vector machines that find optimal decision boundaries, and neural networks that can learn arbitrarily complex mappings. Each has its strengths — linear models are interpretable, trees handle non-linear data naturally, and neural networks excel when data is abundant.
Unsupervised Learning
Unsupervised learning works with unlabeled data. There are no correct answers to learn from — instead, the algorithm must discover structure and patterns on its own. This makes it both more challenging and, in many cases, more powerful than supervised learning, because labeled data is expensive to obtain while unlabeled data is abundant.
The three main tasks in unsupervised learning are clustering, dimensionality reduction, and anomaly detection.
Clustering groups similar data points together. The most popular algorithm is k-means, which partitions data into k clusters by minimizing the distance between each point and its cluster center. The objective function that k-means minimizes is:
Dimensionality reduction compresses high-dimensional data into fewer dimensions while preserving as much information as possible. PCA (Principal Component Analysis) finds the directions of maximum variance and projects data onto them. Autoencoders are neural networks that learn to compress and reconstruct data, discovering efficient representations.
Anomaly detection identifies data points that deviate significantly from normal patterns. This is how credit card fraud detection works — the system learns what normal transactions look like and flags anything unusual. It is also used in manufacturing to catch defective products and in cybersecurity to detect network intrusions.
Reinforcement Learning
Reinforcement learning (RL) is fundamentally different from both supervised and unsupervised learning. Instead of learning from a static dataset, an agent learns by interacting with an environment. At each step, the agent observes the current state, takes an action, receives a reward (or penalty), and transitions to a new state. Over time, the agent learns a policy — a strategy that maximizes cumulative reward.
State → Action → Reward → New State → repeat. The agent's goal is to learn a policy that maximizes the expected cumulative reward over time.
G_t = R_{t+1} + γR_{t+2} + γ²R_{t+3} + …The expected cumulative reward, called the return, is defined as:
Key applications of RL include game playing (AlphaGo, Atari), robotics (learning to walk and grasp objects), autonomous driving, resource management, and even optimizing data center cooling at Google. RL shines in sequential decision-making problems where the optimal strategy is not obvious and must be discovered through exploration.
Beyond the Three Paradigms
The boundaries between learning types are not always rigid. Two important hybrid approaches have emerged:
Semi-supervised learning uses a small amount of labeled data combined with a large amount of unlabeled data. This is practical because labeling data is expensive — a doctor's time to annotate medical images, for instance. The model learns basic structure from the unlabeled data and refines its understanding with the labeled examples. In practice, semi-supervised methods can achieve near-supervised performance with a fraction of the labels.
Self-supervised learning creates its own labels from the data. For example, a language model learns to predict the next word in a sentence — the "label" is just the next word, which comes for free from the text itself. This is how models like GPT and BERT are trained. Similarly, in computer vision, a model might learn to predict the color of a grayscale image or reconstruct a masked patch. Self-supervised learning has become the dominant pretraining strategy in modern AI.
Choosing the Right Approach
Selecting the right type of machine learning depends on your data and your problem:
If you have labeled data and want to predict specific outcomes, use supervised learning. If you have unlabeled data and want to discover hidden patterns, use unsupervised learning. If your problem involves sequential decisions with delayed rewards, reinforcement learning is the way to go. And if you have some labels but not enough, semi-supervised or self-supervised approaches can bridge the gap.
In practice, most real-world ML systems combine multiple paradigms. A recommendation system might use unsupervised clustering to group users, supervised learning to predict ratings, and reinforcement learning to optimize the order of recommendations. Understanding all three paradigms gives you the flexibility to architect effective solutions.
- Supervised learning uses labeled data to predict outcomes — either discrete labels (classification) or continuous values (regression).
- Unsupervised learning discovers hidden structure in unlabeled data through clustering, dimensionality reduction, and anomaly detection.
- Reinforcement learning trains an agent to make sequential decisions by maximizing cumulative reward through trial and error.
- Semi-supervised and self-supervised learning bridge the gap when labeled data is scarce, and modern AI increasingly relies on self-supervised pretraining.
- Choosing the right paradigm depends on your data (labeled vs unlabeled) and your problem type (prediction, discovery, or decision-making).