CNN Architecture
Stacking convolutional layers into a full network requires careful design. How many filters? How deep? When to pool? The classic architectures — LeNet to ResNet — answer these questions and reveal why depth matters so much.
Many Filters in Parallel
A convolutional layer applies F filters simultaneously, producing F feature maps. Each filter learns to detect a different pattern. A layer with 64 filters on a 224×224 input produces 64 feature maps, one per learned detector.
How Many Weights?
Each filter has k×k×C weights plus one bias. A layer with F filters and input channels C has F×(k²C + 1) parameters — independent of spatial size H×W. This is the power of weight sharing.
Conv → Pool → Repeat
The foundational CNN recipe alternates convolutional blocks with pooling layers to progressively shrink spatial size while growing the number of channels. Dense layers at the end map the final feature volume to class scores.
The Pioneer
- Input: 32×32 grayscale (handwritten digits)
- 2 conv + pool stages extract local features
- 3 FC layers classify into 10 digit classes
- ~60K parameters — tiny by modern standards
- Proved that CNNs could outperform fully connected nets on vision tasks
Going Deeper
- AlexNet (2012) — 5 conv layers, ~60M params, won ImageNet by a huge margin; introduced ReLU and dropout
- VGG-16 (2014) — 13 conv layers, all 3×3 filters; proved that depth with small filters beats width with large filters
- Two stacked 3×3 conv layers have the same receptive field as one 5×5 but fewer parameters and more nonlinearities
Skip Connections
Very deep networks suffer from vanishing gradients — signal shrinks to zero before reaching early layers. ResNet solves this with a residual shortcut: the output is F(x) + x. The layer only needs to learn the difference from the identity.
Inception & EfficientNet
- Inception — applies 1×1, 3×3, and 5×5 filters in parallel; concatenates results; lets the network choose its own receptive field per layer
- 1×1 convolutions — mix channels cheaply without changing spatial size; act as learned dimensionality reduction
- EfficientNet — jointly scales depth, width, and resolution via a compound coefficient; state-of-the-art efficiency
Reuse Pretrained Weights
- ImageNet pretrained models have already learned powerful visual features from 1.2 million images
- Feature extraction — freeze all conv layers; train only a new classification head on your data
- Fine-tuning — unfreeze later layers; retrain at a small learning rate to adapt to your domain
- Works even when domains differ: medical imaging, satellite photos, industrial inspection
What You Learned
A conv layer applies F filters, each of size k×k×C, producing F feature maps. The classic Conv → ReLU → Pool pattern shrinks spatial size while growing depth. LeNet pioneered the recipe; AlexNet and VGG proved depth wins; ResNet skip connections made very deep networks trainable. Transfer learning reuses pretrained weights, making powerful CNNs accessible even with small datasets.