CNN Applications
Classification, detection, segmentation, faces, audio, edge deployment — the same convolutional machinery powers all of these. This lesson surveys how CNNs are put to work in the real world.
Softmax & Cross-Entropy
Classification maps an image to a probability over C classes via softmax. Training minimizes cross-entropy loss. Data augmentation and label smoothing improve generalization and calibration at scale.
What & Where?
Detection outputs bounding boxes and class labels. Two-stage detectors (Faster R-CNN) propose regions first, then classify. Quality is measured by IoU between predicted and ground-truth boxes.
Detection in One Pass
- Divide image into S×S grid; each cell predicts B boxes and C class scores simultaneously
- Single forward pass — 45+ FPS, suitable for real-time video
- Non-Maximum Suppression removes overlapping box duplicates
- mAP (mean Average Precision) aggregates precision–recall across all classes
Every Pixel Classified
- Encoder–Decoder — compress with conv+pool, then upsample back to full resolution
- Skip connections pass fine spatial detail from encoder to decoder at each resolution
- FCN (2015) replaced FC layers with 1×1 convolutions for dense predictions
- U-Net concatenates encoder maps to decoder — dominates biomedical segmentation
High-Stakes Perception
- CNNs match or exceed radiologist accuracy on narrow tasks: diabetic retinopathy, pneumonia, skin lesions
- Challenges: small datasets, extreme class imbalance, scanner distribution shift
- 3D U-Net extends to volumetric CT/MRI with 3D convolutions
- Self-supervised pretraining on unlabeled scans overcomes data scarcity
Metric Learning
Standard softmax can’t recognize identities unseen at training time. Metric learning maps faces to embeddings: same person → close, different person → far apart. FaceNet’s triplet loss enforces this geometry.
Content and Style
Freeze the weights and optimize the image instead. Read a layer’s maps directly and you get content — what is present, and where. Read their Gram matrix — channel correlations summed over every position — and position is thrown away, leaving texture and palette. Same activations, two summaries.
Gatys 2016 optimizes per image (minutes). Johnson 2016 trains one feed-forward net per style (one pass). AdaIN 2017 matches feature mean and variance — any style, one pass.
CNNs for Audio & Text
- Spectrograms turn audio into 2D frequency–time images; standard 2D CNNs apply directly
- 1D convolutions slide over raw waveforms or token sequences
- WaveNet — dilated causal 1D conv stacks generate lifelike speech at 16 kHz
- TextCNN — parallel 1D filters of width 2, 3, 4 words for sentence classification
Edge vs. Cloud
- Quantization — INT8 weights cut size 4×, speed up inference on edge hardware
- Pruning — remove near-zero weights or entire channels for sparse networks
- Distillation — train small student on soft outputs of large teacher
- MobileNet — depthwise separable convolutions reduce compute 8–9× vs. standard conv
What You Learned
CNNs underpin image classification, object detection (YOLO, Faster R-CNN), semantic segmentation (U-Net), face recognition (triplet loss), medical imaging, audio (spectrograms + 1D conv), and text. At the edge, quantization, pruning, distillation, and depthwise separable convolutions (MobileNet) compress models to fit tight hardware budgets.