ML 101
M07 · L01
Module 7

Convolution for Images

Images are structured grids of numbers. Convolution exploits that structure — sliding a small filter across the image to detect patterns anywhere, with orders of magnitude fewer parameters than a fully connected layer.

01 / 11
ML 101
M07 · L01
Image Representation

Tensors of Pixels

A grayscale image is a 2D matrix of intensity values 0–255. Color images add a channel dimension: an RGB image of size 224×224 is an H×W×C tensor with 224×224×3 = 150,528 numbers.

Grayscale
H×W×1
Color RGB
H×W×3
02 / 11
ML 101
M07 · L01
The Core Operation

2D Convolution

A filter of size k×k slides across the image. At each position it computes element-wise products with the overlapping patch and sums them into one output value. The same filter is reused everywhere — weight sharing.

Cross-Correlation
(I * K)(i,j) = \sum_{m=0}^{k-1}\sum_{n=0}^{k-1} K(m,n)\cdot I(i+m,j+n)
03 / 11
ML 101
M07 · L01
Filter Intuition

Edges, Blurs, Sharpening

  • Sobel — detects horizontal or vertical edges by differencing neighboring columns or rows
  • Gaussian blur — averages pixels with bell-curve weights, smoothing noise and detail
  • Sharpening — amplifies contrast between each pixel and its neighbors
  • CNNs learn Gabor-like filters automatically — edges in early layers, textures and parts deeper
04 / 11
ML 101
M07 · L01
Stride

Controlling the Step

Stride s makes the filter jump s pixels between positions. Stride 2 halves the output resolution in a single pass — learned downsampling. Modern architectures often use strided convolutions instead of pooling layers.

Output size (no padding)
(H − k) / s + 1
05 / 11
ML 101
M07 · L01
Padding

Preserving Spatial Size

Without padding every convolution shrinks the feature map by k−1. Same padding adds p = (k−1)/2 zeros on each side so the output matches the input. Valid padding means no zeros at all.

Output Size Formula
H_{\text{out}} = \left\lfloor \frac{H - k + 2p}{s} \right\rfloor + 1
06 / 11
ML 101
M07 · L01
Pooling

Spatial Summarization

  • Max pooling — keeps the peak activation in each 2×2 window; retains strongest features
  • Average pooling — takes the mean; retains diffuse information
  • Global average pooling — collapses an entire feature map to one value per channel, replacing a large flatten layer
07 / 11
ML 101
M07 · L01
Receptive Field

How Far Can a Neuron See?

Stacking L layers of k×k convolutions gives each output neuron an effective receptive field of [L(k−1)+1]². Early layers see local edges; deep layers integrate the whole image. Pooling accelerates growth.

Effective Receptive Field
\text{RF} = L(k-1)+1
08 / 11
ML 101
M07 · L01
Why It Works

Three Inductive Biases

  • Local connectivity — meaningful patterns are local; distant pixels share no information
  • Weight sharing — the same edge detector works everywhere in the image
  • Translation equivariance — shifting the input shifts the feature map; the network re-uses detectors across all positions
09 / 11
ML 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

10 / 11
ML 101
Summary
Recap

What You Learned

Images are H×W×C tensors. Convolution slides a k×k filter with weight sharing to detect local patterns. Stride controls step size; padding preserves spatial dimensions. Max pooling summarizes each region. Stacking layers grows the receptive field. Three biases — locality, weight sharing, equivariance — make CNNs uniquely effective for vision.

Next Lesson
11 / 11