ML 101
M11 · L02
Deep Learning in Practice

Transfer Learning & Fine-Tuning

Pretrained model zoos, feature extraction, fine-tuning strategies, LoRA, and few-shot learning — reusing knowledge instead of starting from scratch.

01 / 13
ML 101
M11 · L02
The Core Idea

Don't Start from Scratch

A network trained on 1.2M ImageNet photos already knows edges, textures, and shapes. A model trained on billions of tokens already knows grammar, facts, and reasoning. Reuse these representations for your task — with far less data and compute.

Why It Works
Early layers learn general features (edges, n-grams); later layers learn task-specific abstractions. Only the final layers need significant adaptation.
02 / 13
ML 101
M11 · L02
Model Zoo

The Pretrained Ecosystem

  • Vision CNNs — ResNet, EfficientNet, ConvNeXt (ImageNet pretrained)
  • Vision Transformers — ViT, DeiT, CLIP (patch-based; zero-shot capable)
  • Encoder LMs — BERT, RoBERTa (masked LM; best for classification/NER)
  • Decoder LMs — LLaMA, Mistral, GPT-2 (causal LM; best for generation)
  • Seq2Seq — T5, BART (text-to-text; translation, summarization)
  • Multimodal — BLIP-2, LLaVA (image + text joint representations)
03 / 13
ML 101
M11 · L02
Strategy 1

Feature Extraction

Freeze all pretrained weights (requires_grad = False). Remove the original head. Train only a new head on top of fixed representations.

Fixed Backbone
\hat{y}=W\cdot f_{\text{frozen}}(x)

Best for: small datasets, similar domain. Fast training — often just minutes on CPU.

04 / 13
ML 101
M11 · L02
Strategy 2

Full Fine-Tuning

Unfreeze all weights and update them jointly with the new head using a small learning rate. Use pretrained weights as initialization, not fixed features.

  • Differential LR — smaller rate for early layers, larger for later layers
  • Gradual unfreezing — unfreeze layer groups progressively
  • Catastrophic forgetting — use low LR (1e-5) + early stopping to prevent
  • Warmup schedule: 5–10% of total steps, then cosine decay
05 / 13
ML 101
M11 · L02
Domain Gap

Domain Adaptation

Source and target domains differ — general English vs. medical text, ImageNet vs. satellite images. Pretrained representations are misaligned.

DAPT
Continue pretraining on in-domain unlabeled text
DA
Adversarial alignment or self-training with pseudo-labels
06 / 13
ML 101
M11 · L02
No Gradient Updates

Zero- & Few-Shot Learning

Large models generalize at inference time from task descriptions (zero-shot) or a few examples in the prompt (few-shot in-context learning). No weight updates required.

ICL vs. Fine-Tuning
With >100–1000 labeled examples, fine-tuning beats ICL. With only a prompt, ICL is the only option for API-access models.
07 / 13
ML 101
M11 · L02
LoRA

Low-Rank Adaptation

Freeze W₀. Add trainable low-rank bypass: two small matrices A ∈ ℝ^{d×r} and B ∈ ℝ^{r×k} where r ≪ d,k. Only A and B update.

LoRA Forward
h=W_0 x+\tfrac{\alpha}{r}BAx
08 / 13
ML 101
M11 · L02
LoRA Efficiency

128× Fewer Parameters

r=16
Rank · typical attention layer
128×
Fewer trainable params vs. full FT

QLoRA quantizes the base model to 4-bit NF4, enabling fine-tuning 33B–70B models on a single 24 GB consumer GPU. Merge adapters at inference for zero overhead.

09 / 13
ML 101
M11 · L02
PEFT Methods

Beyond LoRA

  • Adapter layers — bottleneck modules between transformer layers; adds inference latency
  • Prefix tuning — trainable prefix tokens injected into attention K and V sequences
  • Prompt tuning — soft continuous prompt embeddings prepended to input; competitive at scale
  • DoRA — decomposes weight into magnitude + direction; often outperforms LoRA
10 / 13
ML 101
M11 · L02
Practical Workflow

Fine-Tuning Checklist

  • Pick the smallest capable base model
  • Clean and balance your dataset first
  • Feature extraction → LoRA → full FT (escalate as needed)
  • LR: 1e-4 for LoRA adapters, 1e-5–5e-5 for full FT
  • Monitor val loss; stop at 1–5 epochs
  • Merge LoRA adapters before deployment
11 / 13
ML 101
Transfer Learning & Fine-Tuning

Check what transferred

Four questions on fine-tuning, LoRA and in-context learning — retrieval, not recognition.

Question 1 of 0
Score 0/0

12 / 13
ML 101
Key Takeaways
Summary

Key Takeaways

  • Transfer learning reuses pretrained representations — less data, less compute
  • Feature extraction: freeze backbone, train head only — fast and safe for small datasets
  • Full fine-tuning: low LR + gradual unfreezing to avoid catastrophic forgetting
  • LoRA / QLoRA: train low-rank updates only — 128× fewer params, consumer GPU friendly
  • Zero/few-shot ICL: generalize at inference time with no weight updates
  • Domain-adaptive pretraining bridges large domain gaps before task fine-tuning
13 / 13