ML 101
M07 · L03
Module 7

CNN Applications

Classification, detection, segmentation, faces, audio, edge deployment — the same convolutional machinery powers all of these. This lesson surveys how CNNs are put to work in the real world.

01 / 12
ML 101
M07 · L03
Image Classification

Softmax & Cross-Entropy

Classification maps an image to a probability over C classes via softmax. Training minimizes cross-entropy loss. Data augmentation and label smoothing improve generalization and calibration at scale.

Human Top-5
≈ 5% err
ResNet-50
≈ 7.5% err
Best Models
< 2% err
02 / 12
ML 101
M07 · L03
Object Detection

What & Where?

Detection outputs bounding boxes and class labels. Two-stage detectors (Faster R-CNN) propose regions first, then classify. Quality is measured by IoU between predicted and ground-truth boxes.

Intersection over Union
\text{IoU} = \frac{|A \cap B|}{|A \cup B|}
03 / 12
ML 101
M07 · L03
YOLO (2016)

Detection in One Pass

  • Divide image into S×S grid; each cell predicts B boxes and C class scores simultaneously
  • Single forward pass — 45+ FPS, suitable for real-time video
  • Non-Maximum Suppression removes overlapping box duplicates
  • mAP (mean Average Precision) aggregates precision–recall across all classes
04 / 12
ML 101
M07 · L03
Semantic Segmentation

Every Pixel Classified

  • Encoder–Decoder — compress with conv+pool, then upsample back to full resolution
  • Skip connections pass fine spatial detail from encoder to decoder at each resolution
  • FCN (2015) replaced FC layers with 1×1 convolutions for dense predictions
  • U-Net concatenates encoder maps to decoder — dominates biomedical segmentation
05 / 12
ML 101
M07 · L03
Medical Imaging

High-Stakes Perception

  • CNNs match or exceed radiologist accuracy on narrow tasks: diabetic retinopathy, pneumonia, skin lesions
  • Challenges: small datasets, extreme class imbalance, scanner distribution shift
  • 3D U-Net extends to volumetric CT/MRI with 3D convolutions
  • Self-supervised pretraining on unlabeled scans overcomes data scarcity
06 / 12
ML 101
M07 · L03
Face Recognition

Metric Learning

Standard softmax can’t recognize identities unseen at training time. Metric learning maps faces to embeddings: same person → close, different person → far apart. FaceNet’s triplet loss enforces this geometry.

Triplet Loss
\mathcal{L} = \max\!\left(0,\; d(a,p) - d(a,n) + \alpha\right)
07 / 12
ML 101
M07 · L03
Style Transfer

Content and Style

Freeze the weights and optimize the image instead. Read a layer’s maps directly and you get content — what is present, and where. Read their Gram matrix — channel correlations summed over every position — and position is thrown away, leaving texture and palette. Same activations, two summaries.

Gram Matrix
G^{\ell}_{ij} = \sum_{k} F^{\ell}_{ik} F^{\ell}_{jk}

Gatys 2016 optimizes per image (minutes). Johnson 2016 trains one feed-forward net per style (one pass). AdaIN 2017 matches feature mean and variance — any style, one pass.

08 / 12
ML 101
M07 · L03
Beyond Images

CNNs for Audio & Text

  • Spectrograms turn audio into 2D frequency–time images; standard 2D CNNs apply directly
  • 1D convolutions slide over raw waveforms or token sequences
  • WaveNet — dilated causal 1D conv stacks generate lifelike speech at 16 kHz
  • TextCNN — parallel 1D filters of width 2, 3, 4 words for sentence classification
09 / 12
ML 101
M07 · L03
Deployment

Edge vs. Cloud

  • Quantization — INT8 weights cut size 4×, speed up inference on edge hardware
  • Pruning — remove near-zero weights or entire channels for sparse networks
  • Distillation — train small student on soft outputs of large teacher
  • MobileNet — depthwise separable convolutions reduce compute 8–9× vs. standard conv
10 / 12
ML 101
Knowledge Check

Check what stuck

Five questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

11 / 12
ML 101
Summary
Recap

What You Learned

CNNs underpin image classification, object detection (YOLO, Faster R-CNN), semantic segmentation (U-Net), face recognition (triplet loss), medical imaging, audio (spectrograms + 1D conv), and text. At the edge, quantization, pruning, distillation, and depthwise separable convolutions (MobileNet) compress models to fit tight hardware budgets.

Module 7 Complete
Convolutional Neural Networks ✓
12 / 12