Reading
Stories Mode

Types of Machine Learning

~14 min read Lesson 2 of 4 in Module 1

The Three Paradigms

In the previous lesson, we learned that machine learning is about finding patterns in data. But not all learning is the same. Just as humans learn differently — from teachers, from exploration, and from trial and error — machine learning algorithms fall into distinct paradigms, each suited to different kinds of problems.

The three main paradigms are supervised learning, unsupervised learning, and reinforcement learning. Understanding when to use each one is one of the most important skills in applied ML. Let us explore each in depth.

Quick Analogy

Supervised = learning with a teacher who provides the answers. Unsupervised = exploring a library with no catalog and organizing the books yourself. Reinforcement = learning to ride a bike by falling down and getting back up.

Supervised Learning

Supervised learning is the most widely used paradigm in machine learning. The idea is simple: you provide the algorithm with labeled training data — pairs of inputs and their correct outputs — and the model learns a function that maps inputs to outputs. Once trained, it can predict the output for new, unseen inputs.

Supervised learning problems fall into two categories. Classification predicts a discrete label — is this email spam or not? Is this tumor malignant or benign? Regression predicts a continuous value — what will this house sell for? How much revenue will we generate next quarter?

Classification
Discrete Output
Predicts categories: spam/not spam, cat/dog, positive/negative sentiment. Output is one of a fixed set of labels.
Regression
Continuous Output
Predicts numerical values: house prices, temperature forecasts, stock prices. Output is a number on a continuous scale.

The training process works by minimizing a loss function — a measure of how far the model's predictions are from the true labels. For regression, the most common loss function is Mean Squared Error:

Mean Squared Error (MSE)
J(\theta) = \frac{1}{n}\sum_{i=1}^{n}\left(h_\theta(x^{(i)}) - y^{(i)}\right)^2
The MSE loss measures the average squared difference between predictions and true values — penalizing large errors more heavily

Common supervised learning algorithms include linear regression for predicting continuous values, decision trees that split data based on feature thresholds, support vector machines that find optimal decision boundaries, and neural networks that can learn arbitrarily complex mappings. Each has its strengths — linear models are interpretable, trees handle non-linear data naturally, and neural networks excel when data is abundant.

Unsupervised Learning

Unsupervised learning works with unlabeled data. There are no correct answers to learn from — instead, the algorithm must discover structure and patterns on its own. This makes it both more challenging and, in many cases, more powerful than supervised learning, because labeled data is expensive to obtain while unlabeled data is abundant.

The three main tasks in unsupervised learning are clustering, dimensionality reduction, and anomaly detection.

Clustering groups similar data points together. The most popular algorithm is k-means, which partitions data into k clusters by minimizing the distance between each point and its cluster center. The objective function that k-means minimizes is:

K-Means Objective
J = \sum_{k=1}^{K}\sum_{x_i \in C_k} \|x_i - \mu_k\|^2
K-means minimizes the sum of squared distances between each data point and its assigned cluster centroid

Dimensionality reduction compresses high-dimensional data into fewer dimensions while preserving as much information as possible. PCA (Principal Component Analysis) finds the directions of maximum variance and projects data onto them. Autoencoders are neural networks that learn to compress and reconstruct data, discovering efficient representations.

Anomaly detection identifies data points that deviate significantly from normal patterns. This is how credit card fraud detection works — the system learns what normal transactions look like and flags anything unusual. It is also used in manufacturing to catch defective products and in cybersecurity to detect network intrusions.

Reinforcement Learning

Reinforcement learning (RL) is fundamentally different from both supervised and unsupervised learning. Instead of learning from a static dataset, an agent learns by interacting with an environment. At each step, the agent observes the current state, takes an action, receives a reward (or penalty), and transitions to a new state. Over time, the agent learns a policy — a strategy that maximizes cumulative reward.

The RL Loop

State → Action → Reward → New State → repeat. The agent's goal is to learn a policy that maximizes the expected cumulative reward over time.

G_t = R_{t+1} + γR_{t+2} + γ²R_{t+3} + …

The expected cumulative reward, called the return, is defined as:

Expected Return
G_t = \sum_{k=0}^{\infty} \gamma^k R_{t+k+1}, \quad 0 \leq \gamma < 1
The discount factor γ (between 0 and 1) controls how much the agent values future rewards versus immediate ones

Key applications of RL include game playing (AlphaGo, Atari), robotics (learning to walk and grasp objects), autonomous driving, resource management, and even optimizing data center cooling at Google. RL shines in sequential decision-making problems where the optimal strategy is not obvious and must be discovered through exploration.

Beyond the Three Paradigms

The boundaries between learning types are not always rigid. Two important hybrid approaches have emerged:

Semi-supervised learning uses a small amount of labeled data combined with a large amount of unlabeled data. This is practical because labeling data is expensive — a doctor's time to annotate medical images, for instance. The model learns basic structure from the unlabeled data and refines its understanding with the labeled examples. In practice, semi-supervised methods can achieve near-supervised performance with a fraction of the labels.

Self-supervised learning creates its own labels from the data. For example, a language model learns to predict the next word in a sentence — the "label" is just the next word, which comes for free from the text itself. This is how models like GPT and BERT are trained. Similarly, in computer vision, a model might learn to predict the color of a grayscale image or reconstruct a masked patch. Self-supervised learning has become the dominant pretraining strategy in modern AI.

Choosing the Right Approach

Selecting the right type of machine learning depends on your data and your problem:

Decision Guide
?
Have Labels?
Yes → Supervised
?
Find Structure?
Yes → Unsupervised
?
Sequential?
Yes → Reinforcement
?
Few Labels?
Yes → Semi-supervised

If you have labeled data and want to predict specific outcomes, use supervised learning. If you have unlabeled data and want to discover hidden patterns, use unsupervised learning. If your problem involves sequential decisions with delayed rewards, reinforcement learning is the way to go. And if you have some labels but not enough, semi-supervised or self-supervised approaches can bridge the gap.

In practice, most real-world ML systems combine multiple paradigms. A recommendation system might use unsupervised clustering to group users, supervised learning to predict ratings, and reinforcement learning to optimize the order of recommendations. Understanding all three paradigms gives you the flexibility to architect effective solutions.

In the next two lessons, we will explore supervised learning in much greater depth — first building a model and walking through the training loop, then learning how to judge whether it actually works, with metrics like accuracy, precision, and recall.

Key Takeaways
  • Supervised learning uses labeled data to predict outcomes — either discrete labels (classification) or continuous values (regression).
  • Unsupervised learning discovers hidden structure in unlabeled data through clustering, dimensionality reduction, and anomaly detection.
  • Reinforcement learning trains an agent to make sequential decisions by maximizing cumulative reward through trial and error.
  • Semi-supervised and self-supervised learning bridge the gap when labeled data is scarce, and modern AI increasingly relies on self-supervised pretraining.
  • Choosing the right paradigm depends on your data (labeled vs unlabeled) and your problem type (prediction, discovery, or decision-making).
Previous What is ML? Module Overview Next Lesson Your First ML Model