Transfer Learning & Fine-Tuning
Pretrained model zoos, feature extraction, fine-tuning strategies, LoRA, and few-shot learning — reusing knowledge instead of starting from scratch.
Don't Start from Scratch
A network trained on 1.2M ImageNet photos already knows edges, textures, and shapes. A model trained on billions of tokens already knows grammar, facts, and reasoning. Reuse these representations for your task — with far less data and compute.
The Pretrained Ecosystem
- Vision CNNs — ResNet, EfficientNet, ConvNeXt (ImageNet pretrained)
- Vision Transformers — ViT, DeiT, CLIP (patch-based; zero-shot capable)
- Encoder LMs — BERT, RoBERTa (masked LM; best for classification/NER)
- Decoder LMs — LLaMA, Mistral, GPT-2 (causal LM; best for generation)
- Seq2Seq — T5, BART (text-to-text; translation, summarization)
- Multimodal — BLIP-2, LLaVA (image + text joint representations)
Feature Extraction
Freeze all pretrained weights (requires_grad = False). Remove the original head. Train only a new head on top of fixed representations.
Best for: small datasets, similar domain. Fast training — often just minutes on CPU.
Full Fine-Tuning
Unfreeze all weights and update them jointly with the new head using a small learning rate. Use pretrained weights as initialization, not fixed features.
- Differential LR — smaller rate for early layers, larger for later layers
- Gradual unfreezing — unfreeze layer groups progressively
- Catastrophic forgetting — use low LR (1e-5) + early stopping to prevent
- Warmup schedule: 5–10% of total steps, then cosine decay
Domain Adaptation
Source and target domains differ — general English vs. medical text, ImageNet vs. satellite images. Pretrained representations are misaligned.
Zero- & Few-Shot Learning
Large models generalize at inference time from task descriptions (zero-shot) or a few examples in the prompt (few-shot in-context learning). No weight updates required.
Low-Rank Adaptation
Freeze W₀. Add trainable low-rank bypass: two small matrices A ∈ ℝ^{d×r} and B ∈ ℝ^{r×k} where r ≪ d,k. Only A and B update.
128× Fewer Parameters
QLoRA quantizes the base model to 4-bit NF4, enabling fine-tuning 33B–70B models on a single 24 GB consumer GPU. Merge adapters at inference for zero overhead.
Beyond LoRA
- Adapter layers — bottleneck modules between transformer layers; adds inference latency
- Prefix tuning — trainable prefix tokens injected into attention K and V sequences
- Prompt tuning — soft continuous prompt embeddings prepended to input; competitive at scale
- DoRA — decomposes weight into magnitude + direction; often outperforms LoRA
Fine-Tuning Checklist
- Pick the smallest capable base model
- Clean and balance your dataset first
- Feature extraction → LoRA → full FT (escalate as needed)
- LR: 1e-4 for LoRA adapters, 1e-5–5e-5 for full FT
- Monitor val loss; stop at 1–5 epochs
- Merge LoRA adapters before deployment
Key Takeaways
- Transfer learning reuses pretrained representations — less data, less compute
- Feature extraction: freeze backbone, train head only — fast and safe for small datasets
- Full fine-tuning: low LR + gradual unfreezing to avoid catastrophic forgetting
- LoRA / QLoRA: train low-rank updates only — 128× fewer params, consumer GPU friendly
- Zero/few-shot ICL: generalize at inference time with no weight updates
- Domain-adaptive pretraining bridges large domain gaps before task fine-tuning