Convolution for Images
Images are structured grids of numbers. Convolution exploits that structure — sliding a small filter across the image to detect patterns anywhere, with orders of magnitude fewer parameters than a fully connected layer.
Tensors of Pixels
A grayscale image is a 2D matrix of intensity values 0–255. Color images add a channel dimension: an RGB image of size 224×224 is an H×W×C tensor with 224×224×3 = 150,528 numbers.
2D Convolution
A filter of size k×k slides across the image. At each position it computes element-wise products with the overlapping patch and sums them into one output value. The same filter is reused everywhere — weight sharing.
Edges, Blurs, Sharpening
- Sobel — detects horizontal or vertical edges by differencing neighboring columns or rows
- Gaussian blur — averages pixels with bell-curve weights, smoothing noise and detail
- Sharpening — amplifies contrast between each pixel and its neighbors
- CNNs learn Gabor-like filters automatically — edges in early layers, textures and parts deeper
Controlling the Step
Stride s makes the filter jump s pixels between positions. Stride 2 halves the output resolution in a single pass — learned downsampling. Modern architectures often use strided convolutions instead of pooling layers.
Preserving Spatial Size
Without padding every convolution shrinks the feature map by k−1. Same padding adds p = (k−1)/2 zeros on each side so the output matches the input. Valid padding means no zeros at all.
Spatial Summarization
- Max pooling — keeps the peak activation in each 2×2 window; retains strongest features
- Average pooling — takes the mean; retains diffuse information
- Global average pooling — collapses an entire feature map to one value per channel, replacing a large flatten layer
How Far Can a Neuron See?
Stacking L layers of k×k convolutions gives each output neuron an effective receptive field of [L(k−1)+1]². Early layers see local edges; deep layers integrate the whole image. Pooling accelerates growth.
Three Inductive Biases
- Local connectivity — meaningful patterns are local; distant pixels share no information
- Weight sharing — the same edge detector works everywhere in the image
- Translation equivariance — shifting the input shifts the feature map; the network re-uses detectors across all positions
What You Learned
Images are H×W×C tensors. Convolution slides a k×k filter with weight sharing to detect local patterns. Stride controls step size; padding preserves spatial dimensions. Max pooling summarizes each region. Stacking layers grows the receptive field. Three biases — locality, weight sharing, equivariance — make CNNs uniquely effective for vision.