Images as Structured Grids
A grayscale image is simply a 2D matrix of numbers. Each entry — a pixel — holds an intensity value between 0 (black) and 255 (white). A 28×28 MNIST digit is a matrix with 784 numbers. A 1080p photograph has 1,920×1,080 = 2,073,600 pixels.
Color images add a third dimension: channels. A standard RGB image has three channels — red, green, and blue — stacked into a tensor of shape H×W×3. So a 224×224 RGB image is a 3D tensor with 224×224×3 = 150,528 numbers. Some domains use more channels: medical CT scans can have dozens of slices; satellite imagery can have hyperspectral channels beyond the visible range.
What makes images special, and why does this structure matter? Pixel values are not independent. Nearby pixels tend to be similar (spatial locality), and the same pattern — an edge, a corner, a texture — can appear anywhere in the image (translation invariance). A fully connected layer that treats every pixel independently ignores both properties and scales catastrophically: a single dense layer connecting 150,528 inputs to just 1,000 neurons already has 150 million parameters. Convolutional neural networks exploit image structure to be dramatically more efficient and effective.
The Convolution Operation
The core operation is a convolution (more precisely, a cross-correlation in most deep learning implementations). A small matrix called a filter or kernel slides across the input image. At each position, it computes an element-wise product between the filter weights and the overlapping patch of the image, then sums all products into a single output value. The result at every position forms an output feature map.
The same filter is applied at every spatial position. This weight sharing is the key efficiency gain: a 3×3 filter has only 9 learnable parameters regardless of image size, and it learns to detect the same local pattern everywhere in the image. Compare this to a fully connected layer where every output neuron has its own set of weights for every input pixel.
Mathematically, true convolution flips the filter before sliding it (180° rotation). Most deep learning libraries implement cross-correlation, which skips the flip, and call it “convolution” anyway. Since the filter weights are learned, the distinction is irrelevant: the network learns whatever flipped or unflipped weights it needs.
What Filters Detect: Edges, Blurs, Sharpening
Before CNNs, computer vision engineers designed filters by hand. The convolution operation is the same — only the values inside the filter differ. Understanding these hand-crafted filters builds intuition for what learned filters discover automatically.
The Sobel filter detects horizontal or vertical edges. The horizontal Sobel kernel is [−1, 0, 1; −2, 0, 2; −1, 0, 1]: it subtracts the left column from the right, amplifying regions where intensity changes rapidly left to right. Applied to an image of a face, it highlights the edges of eyes, nose, and mouth as bright lines against a dark background.
A Gaussian blur filter averages nearby pixels with Gaussian weights, smoothing the image. This reduces noise and high-frequency detail. The filter values form a bell curve: the central pixel gets the most weight, falling off toward the edges. In CNNs, early layers often learn Gabor-like filters that combine edge detection with orientation selectivity — similar to cells in the mammalian visual cortex.
A sharpening filter does the opposite: it amplifies the difference between each pixel and its neighbors, making edges more pronounced. A typical sharpening kernel has a large positive center and negative surroundings: [0, −1, 0; −1, 5, −1; 0, −1, 0].
In the first layer of a trained CNN, filters often resemble Gabor functions — oriented edge and color detectors. Deeper layers combine these into textures, then parts, then high-level concepts. This hierarchical composition — edges to textures to parts to objects — emerges automatically from gradient descent, mirroring the ventral stream of the visual cortex.
Stride: Controlling the Step Size
By default, the filter moves one pixel at a time (stride 1). Increasing the stride to s means the filter jumps s pixels between positions. A stride of 2 on a 6×6 input with a 3×3 filter (no padding) produces a 2×2 output instead of a 4×4 output — the spatial resolution is halved.
Stride is one way to downsample the spatial dimensions. A strided convolution computes the same operation as a regular convolution followed by picking every s-th output — but it does so in a single pass. Strided convolutions with stride 2 are commonly used in modern architectures to replace pooling layers, offering learned downsampling rather than a fixed operation.
Padding: Preserving Spatial Size
Without any padding, each convolution shrinks the spatial dimensions. For a k×k filter, the output is (H−k+1)×(W−k+1). Stack many layers without padding and the feature map rapidly shrinks to nothing.
Zero-padding adds a border of zeros around the input before applying the filter. Same padding adds enough zeros so the output has the same spatial size as the input (for stride 1). For a k×k filter, this requires padding p = (k−1)/2 on each side. Valid padding means no padding at all — the filter is only applied where it fully overlaps the input.
Pooling: Spatial Summarization
After a convolutional layer, the feature maps are often downsampled using a pooling operation. Pooling reduces spatial resolution while retaining the strongest detected features, making the representation more compact and providing a degree of translation invariance.
Max pooling divides the feature map into non-overlapping windows (typically 2×2) and keeps the maximum value in each. The intuition: if a feature was detected anywhere in that 2×2 region, keep the strongest detection. Max pooling is the most common choice because it selects the most active response, discarding weak activations.
Average pooling takes the mean of each window instead of the maximum. It retains diffuse information rather than peak values. Average pooling is used at the end of many modern architectures (global average pooling) to collapse an entire feature map into a single number per channel before the final classification layer.
Before the final dense layer, older architectures (AlexNet, VGG) flatten the feature maps into a long vector, leading to millions of parameters in the final dense layer. Modern architectures use global average pooling instead: reduce each channel to its spatial average, producing a vector of length equal to the number of channels. This drastically cuts parameters and adds spatial invariance. A 7×7×512 feature map becomes a 512-dimensional vector with global average pooling but a 25,088-dimensional one with flatten.
Receptive Field and Hierarchical Features
The receptive field of a neuron is the region of the input image that can influence its activation. A single 3×3 convolutional layer has a receptive field of 3×3. Two stacked 3×3 layers have an effective receptive field of 5×5. Three layers give a 7×7 receptive field. In general, stacking L layers of k×k convolutions gives an effective receptive field of [L(k−1)+1]×[L(k−1)+1].
This is the key to hierarchical feature learning. Early layers have small receptive fields and detect local patterns: edges, corners, color blobs. Middle layers combine these into larger structures: textures, object parts. Deep layers integrate globally to recognize entire objects. Each layer sees a larger context than the one below it.
Pooling accelerates the growth of the effective receptive field. After a 2×2 max pool, each subsequent convolutional position covers twice the original spatial extent. A network alternating convolutional and pooling layers sees the entire image in relatively few layers — the famous architectures AlexNet and VGG achieve this in 5–8 convolutional stages.
Why Convolution Works for Images
Three inductive biases make convolution uniquely suited to visual data:
Local connectivity. Each output neuron connects only to a small patch of the input, not the entire image. This is justified because meaningful patterns — edges, textures — are local: a pixel at position (10, 10) tells you nothing about a distant pixel at (200, 200) when detecting a local edge.
Weight sharing. The same filter is applied at every spatial position. An edge detector learned on one part of the image should work equally well on any other part. This assumption — that image statistics are approximately stationary across space — is broadly true and reduces parameters by a factor equal to the number of positions.
Translation equivariance. If a cat appears in the top-left of one image and the bottom-right of another, the convolutional layers produce the same activations, just shifted in space. This means the network does not need to re-learn cat detection for every possible position — a dramatic simplification compared to fully connected networks.
Convolutional layers are equivariant to translation: shifting the input shifts the feature map by the same amount. Pooling introduces invariance: the output is unchanged by small translations. Classification heads need invariance (the label “cat” should not change with position), so CNNs combine equivariant convolutions with invariant pooling to get the best of both. Note that convolutions are not inherently invariant to rotation or scale — these must be handled through data augmentation or specialized architectures.
- Images are H×W×C tensors. Grayscale has C=1; RGB has C=3. Pixel values encode intensity or color, and nearby pixels are spatially correlated — a property that CNNs exploit.
- Convolution slides a small filter over the input, computing element-wise products and summing them at each position. Weight sharing gives a k×k filter only k² parameters regardless of image size.
- Hand-crafted filters (Sobel, Gaussian, sharpening) reveal what convolution can detect. CNNs learn filters automatically, discovering edges in early layers and complex features in deeper layers.
- Stride controls how many pixels the filter jumps per step. Larger stride reduces output size. Padding (same vs. valid) controls whether the output spatial size is preserved or shrinks.
- Max pooling takes the maximum in each spatial window; average pooling takes the mean. Both downsample the feature map. Global average pooling collapses an entire spatial map to a single number per channel.
- The receptive field of a neuron is the input region it can see. Stacking L layers of k×k convolutions gives an effective receptive field of [L(k−1)+1]². Pooling accelerates receptive field growth.
- Convolution works because of three inductive biases: local connectivity (patterns are local), weight sharing (the same pattern appears everywhere), and translation equivariance (shifted input, shifted output). Pooling adds invariance.