Reading
Stories Mode

CNN Architecture

~24 min read Lesson 2 of 3 in Module 7

From Operations to Architecture

The previous lesson introduced the building blocks: convolution, pooling, stride, and padding. Now we assemble those blocks into full architectures. A CNN architecture is the specific sequence and configuration of layers — how many convolutional layers, how many filters each has, where pooling happens, and how the spatial feature maps eventually connect to output predictions.

Architecture design determines almost everything: the number of parameters, the size of the receptive field, how quickly gradients flow during training, and ultimately how well the network generalizes. The history of CNNs is largely a history of better architectures — each breakthrough solving a specific failure mode of the previous generation.

The Convolutional Layer in Detail

A single convolutional layer is defined by four hyperparameters: the number of filters F, the filter size k×k, the stride s, and the padding p. Each filter produces one output feature map — a 2D map of responses to that filter’s pattern across all input positions. With F filters, the layer outputs F feature maps stacked into a 3D tensor.

For a color image, each filter has depth equal to the number of input channels C. A 3×3 filter applied to an RGB image is actually a 3×3×3 = 27-parameter tensor. The filter slides across the spatial dimensions (height and width) while spanning the full channel depth at each position. The output feature map is a single 2D map per filter — the channel dimension is collapsed by the summation inside convolution.

Parameters per Convolutional Layer
\text{Params} = F \times (k^2 \times C + 1)
A convolutional layer with F filters, each of size k×k×C (where C is the number of input channels), has F×(k²×C + 1) parameters — the +1 accounts for the bias term per filter. Compare to a fully connected layer mapping C×H×W inputs to F outputs: that would require F×(C×H×W + 1) parameters, orders of magnitude more for large images.

After each convolutional layer, a nonlinear activation function is applied element-wise to the feature maps. The dominant choice is ReLU (Rectified Linear Unit): f(x) = max(0, x). ReLU sets negative activations to zero and keeps positive ones unchanged. It is computationally cheap, avoids the vanishing gradient problem that plagues sigmoid and tanh for deep networks, and works remarkably well in practice.

Why ReLU Over Sigmoid?

Sigmoid squashes inputs into (0, 1) and saturates for large positive or negative values — gradients become nearly zero there, slowing learning. ReLU has gradient exactly 1 for positive inputs, so gradients flow freely through many layers. The cost is “dead neurons”: units that always output 0 and never recover. Leaky ReLU and ELU address this by allowing small non-zero gradients for negative inputs.

The Classic Conv–Pool Pattern

The standard blueprint, established by LeNet and later AlexNet, alternates convolutional stages with pooling stages: Conv → ReLU → Pool → Conv → ReLU → Pool → … → Flatten → FC → Softmax. Each convolutional stage detects features at progressively larger scales. Each pooling step halves the spatial resolution, growing the effective receptive field.

As spatial dimensions shrink, the number of filters typically doubles. A common progression: 32 filters in layer 1, 64 in layer 2, 128 in layer 3. This keeps the total number of activations roughly constant across layers (halved spatial area compensates for doubled channel count) and progressively enriches the feature representation.

After the final pooling layer, the 3D feature tensor is flattened into a 1D vector and fed into one or more fully connected (FC) layers. The final FC layer outputs one logit per class; a softmax converts logits to a probability distribution. For C classes, the softmax output is: p_i = exp(z_i) / Σ_j exp(z_j), where z_i is the i-th logit.

LeNet to AlexNet: Proof of Concept

LeNet-5 (LeCun et al., 1998) was among the first successful CNNs, designed for handwritten digit recognition on MNIST. It had two convolutional layers (5×5 filters, 6 and 16 feature maps), average pooling, and three fully connected layers — about 60,000 parameters in total. LeNet showed CNNs could learn hierarchical features from raw pixels, but remained limited to small, low-resolution inputs.

AlexNet (Krizhevsky, Sutskever, Hinton, 2012) was the architecture that reignited the field. Winning the ImageNet LSVRC competition by a dramatic margin (15.3% top-5 error vs. 26.2% for the runner-up), it demonstrated that deep CNNs could dominate large-scale image recognition. AlexNet had five convolutional layers, three FC layers, ~60 million parameters, and critically, it used ReLU instead of sigmoid, dropout for regularization, and GPUs for training. Larger input (224×224) and data augmentation were also key.

The ImageNet Moment

The 2012 ImageNet competition was a watershed moment in AI history. AlexNet’s 10.8 percentage-point improvement over the previous state of the art shocked the field. Within two years, every top-performing team used CNNs. The GPU-accelerated training approach AlexNet pioneered is now the universal standard for deep learning.

VGGNet (Simonyan and Zisserman, 2014) explored the effect of depth with a very simple recipe: use only 3×3 convolutions (the smallest filter that can still capture left/right and up/down spatial relationships), always same-padding, and double the number of filters after each max pool. VGG-16 has 16 weight layers and ~138 million parameters. Its regularity made it easy to understand and adapt, and VGG features became the standard off-the-shelf representations for years.

Two 3×3 Layers Beat One 5×5

VGG’s key insight: two stacked 3×3 convolutional layers have the same effective receptive field as a single 5×5 layer, but with fewer parameters (2×9 = 18 vs. 25 per channel) and an extra nonlinearity in between, increasing discriminative power. Three stacked 3×3 layers match a 7×7 layer with even greater efficiency. This principle — prefer many small filters over few large ones — became a lasting design rule.

ResNet: Solving the Depth Problem

Simply stacking more layers does not always help. In the early 2010s, practitioners observed that very deep networks (20+ layers) performed worse than shallower ones on training data — not just due to overfitting, but because of an optimization failure. Gradients vanished or exploded as they propagated back through dozens of layers, and the network struggled to learn even the identity mapping.

ResNet (He et al., 2015) solved this with residual connections (also called skip connections). Instead of learning a mapping H(x) directly, each residual block learns the residual F(x) = H(x) − x, then adds the input x back: output = F(x) + x. This seemingly minor change has a profound effect: gradients can flow directly through the skip connection, bypassing the nonlinear layers entirely, enabling training of networks with hundreds or even thousands of layers.

Residual Block
y = F(x,\{W_i\}) + x
A residual block adds the input x directly to the output of two (or three) convolutional layers F(x, {W_i}). The gradient of the loss with respect to x receives contributions from both the residual branch and the skip connection, so learning is stable even at extreme depth. When dimensions change (e.g., due to strided convolution), a 1×1 convolution W_s projects x to the new size: y = F(x, {W_i}) + W_s x.

ResNet-50 achieved 3.6% top-5 error on ImageNet — surpassing human-level performance for the first time. Its family spans ResNet-18 (11M params) to ResNet-152 (60M params). The bottleneck design in deeper variants uses 1×1 → 3×3 → 1×1 convolutions to reduce and then restore channel depth, keeping computation tractable.

Inception and EfficientNet: Width and Scaling

GoogLeNet / Inception (Szegedy et al., 2014) took a different approach: rather than choosing a single filter size, use multiple in parallel. The Inception module applies 1×1, 3×3, and 5×5 convolutions simultaneously, plus a max pool branch, then concatenates their outputs along the channel dimension. This lets the network decide which scale is most useful for each feature. 1×1 convolutions before the larger filters act as a bottleneck, dramatically reducing computation.

EfficientNet (Tan and Le, 2019) introduced a principled approach to scaling CNNs. Previous models scaled only one dimension at a time: VGG scaled depth, WideResNet scaled width, and some models scaled input resolution. EfficientNet showed that compound scaling — jointly scaling depth, width, and resolution with a fixed ratio — consistently yields better accuracy for the same computational budget. A single scaling coefficient φ controls how aggressively all three dimensions grow together.

Architecture Benchmarks (ImageNet Top-1)

LeNet ≈ 99% on MNIST digits (1998) • AlexNet ≈ 63% on ImageNet (2012) • VGG-16 ≈ 74% (2014) • ResNet-50 ≈ 76% (2015) • EfficientNet-B7 ≈ 84% (2019). Each generation roughly halved the error rate of the previous, often while also reducing parameter count. Today, vision transformers push further, but CNNs remain widely deployed for their efficiency.

Transfer Learning: Stand on Giants’ Shoulders

Training a CNN from scratch on ImageNet requires millions of labeled images, weeks of GPU time, and careful hyperparameter tuning. For most practical problems, this is unnecessary: a network already trained on ImageNet has learned rich, general visual features — edges, textures, shapes, object parts — that transfer remarkably well to new domains.

Feature extraction is the simplest transfer strategy: take a pretrained network, remove its final classification layer, freeze all weights, and use the remaining network as a fixed feature extractor. Run your images through the frozen network, collect the feature vectors from the penultimate layer, and train a simple classifier (logistic regression, SVM) on those features. This works surprisingly well even when the target domain differs significantly from ImageNet.

Fine-tuning goes further: after connecting a new classification head to the pretrained backbone, continue training some or all of the network on the target dataset with a small learning rate. Typically the final few layers are fine-tuned first (they are most dataset-specific), then training is extended to earlier layers if the target dataset is large enough. Fine-tuning generally outperforms feature extraction when sufficient labeled data is available.

When to Freeze vs. Fine-Tune

A useful heuristic: if your target dataset is small and similar to ImageNet, freeze everything and train only the head — more parameters than data leads to overfitting. If your dataset is large and similar, fine-tune the whole network. If it is small but very different (e.g., medical X-rays), fine-tune only the top layers carefully with aggressive regularization. If it is large and different, full fine-tuning is safe and usually best.

Transfer learning has democratized computer vision. Researchers with a few hundred labeled images and a laptop GPU can reach competitive accuracy on many tasks by fine-tuning a pretrained ResNet or EfficientNet. Hubs like PyTorch Image Models (timm) and TensorFlow Hub provide hundreds of pretrained checkpoints ready to adapt.

Key Takeaways
  • A convolutional layer with F filters of size k×k×C has F×(k²×C + 1) parameters. ReLU activation follows each conv layer; it avoids vanishing gradients and is computationally cheap.
  • The classic pattern is Conv → ReLU → Pool, repeated, then Flatten → FC → Softmax. Filters double as spatial dimensions halve, keeping activation volume roughly constant.
  • LeNet proved the concept on small images. AlexNet scaled it to ImageNet with GPUs, ReLU, and dropout. VGG showed that depth with uniform 3×3 filters is highly effective; two 3×3 layers match one 5×5 at lower cost.
  • ResNet solved the degradation problem with residual (skip) connections: y = F(x) + x. Gradients flow through the skip path, enabling training of hundreds of layers stably.
  • Inception uses parallel multi-scale filters and concatenates their outputs. EfficientNet applies compound scaling to jointly grow depth, width, and resolution for maximum efficiency.
  • Transfer learning: use a network pretrained on ImageNet as a feature extractor or fine-tune it on your dataset. Feature extraction works with small datasets; fine-tuning excels when more data is available. Freeze early layers; unfreeze later layers progressively.
Previous Convolution for Images Overview Next CNN Applications