Home / ML 101 / Module 11 / Lesson 2

Transfer Learning & Fine-Tuning

Pretrained model zoos, feature extraction, fine-tuning strategies, domain adaptation, few-shot learning, and LoRA — reusing knowledge instead of starting from scratch.

~18 min read M11 · L2 Advanced

The Core Idea: Don't Start from Scratch

Training a deep neural network from random weights on a small dataset is a recipe for overfitting. The model has millions of parameters but only thousands of examples — it memorizes the training set rather than learning general features. Transfer learning solves this by reusing weights learned on a large, related dataset as a starting point for a new task.

The intuition is compelling: a convolutional network trained on 1.2 million ImageNet photos has already learned to detect edges, textures, patterns, and object parts. A network trained on billions of text tokens has encoded grammar, facts, and reasoning patterns. These generic representations transfer surprisingly well to specific downstream tasks — medical imaging, legal document classification, code generation — with far less data and compute than training from scratch.

Why Transfer Works

Early layers in deep networks learn low-level, general features (edges, n-grams, phonemes); later layers learn task-specific abstractions. When source and target tasks share similar input distributions, early-layer features are directly reusable. Only the task-specific head and the final layers need significant adaptation.

The Pretrained Model Zoo

A rich ecosystem of open pretrained models is available through Hugging Face Hub, PyTorch Hub, TensorFlow Hub, and model-specific repositories. The main families:

Vision (CNN)
ImageNet pretrained
ResNet, EfficientNet, ConvNeXt. Trained on ImageNet-1K or 21K. Standard baseline for image classification and feature extraction.
Vision (ViT)
Patch transformer
Vision Transformers (ViT, DeiT, CLIP). Scale better than CNNs; CLIP enables zero-shot image classification via text-image alignment.
Language (Encoder)
BERT-style MLM
BERT, RoBERTa, DeBERTa. Pretrained via masked language modeling. Excellent for classification, NER, and QA tasks.
Language (Decoder)
GPT-style CLM
GPT-2, LLaMA, Mistral. Pretrained via causal language modeling. Excellent for generation, summarization, instruction following.
Encoder-Decoder
Seq2Seq
T5, BART, mBART. Pretrained for text-to-text tasks. Excels at translation, summarization, and structured prediction.
Multimodal
Image + Text
CLIP, BLIP-2, LLaVA. Joint vision-language representations. Enable visual question answering and image captioning.

Feature Extraction: Frozen Backbone

The simplest transfer learning strategy is feature extraction: freeze all pretrained weights (set requires_grad = False), remove the original classification head, and train only a new head on top of the fixed representations. The pretrained backbone acts as a fixed feature extractor; the only learnable parameters are in the new head.

Feature Extraction
\hat{y} = W \cdot f_{\text{frozen}}(x)
f(·) is the frozen pretrained backbone; W is the trainable new head. Only W is updated during training — drastically reducing the number of learnable parameters.

Feature extraction is ideal when: (1) your target dataset is small (hundreds to a few thousand examples), (2) the target domain is close to the pretraining domain, and (3) compute budget is tight. Training only the head is fast — often just minutes on a CPU. The risk is that the frozen representations may not be well-adapted to the target task if domains differ significantly.

Fine-Tuning: Unfreeze and Train

Fine-tuning goes further: after adding a new head, all (or most) weights in the pretrained model are unfrozen and updated jointly with the new head using a small learning rate. The pretrained representations are used as initialization rather than fixed features, allowing them to adapt to the target task.

Learning Rate Strategy

A critical insight: early layers need far less adaptation than later layers, because early features are more general. Differential learning rates (also called discriminative fine-tuning) address this by assigning different learning rates to different layer groups — smaller for early layers, larger for later layers and the new head:

Differential LR
\eta_{L_1} < \eta_{L_2} < \cdots < \eta_{L_k}
Layer groups L1 (early) through Lk (new head) receive progressively larger learning rates. A common factor is 2–10× between adjacent groups.

Gradual Unfreezing

Rather than unfreezing all layers at once, gradual unfreezing starts by training only the new head for a few epochs, then progressively unfreezes earlier layer groups one by one. This prevents the pretrained weights from being destroyed early in training when the random head is producing large gradients.

Catastrophic Forgetting

A well-known failure mode is catastrophic forgetting: fine-tuning on a small target dataset with a large learning rate can overwrite the general representations learned during pretraining, causing performance to collapse. Mitigations include: use a small learning rate (1e-5 to 1e-4), apply early stopping, use dropout, and keep training epochs modest.

StrategyTrainable ParamsBest WhenRisk
Feature extractionHead only (~1%)Small dataset, similar domainSuboptimal if domain differs
Full fine-tuningAll (100%)Moderate dataset, any domainCatastrophic forgetting; expensive
Gradual unfreezingIncreases over timeSmall-to-moderate datasetMore complex training loop
LoRA / PEFT0.1–2%Large models, limited memorySlightly lower peak performance

Domain Adaptation

Domain adaptation addresses the case where the target domain differs significantly from the pretraining domain — for example, using a model pretrained on general English text for medical or legal documents, or using an ImageNet model for satellite imagery. The distributional shift means that pretraining representations are misaligned with the target inputs.

Continued Pretraining

One effective strategy is continued pretraining (also called domain-adaptive pretraining, DAPT): run the pretraining objective (masked language modeling for BERT, causal language modeling for GPT) on a large corpus of in-domain text before fine-tuning on the target task. This aligns the vocabulary statistics and contextual representations with the target domain without requiring labeled data.

Examples: BioBERT and ClinicalBERT were created by continuing BERT pretraining on PubMed and clinical notes respectively, then fine-tuning on biomedical NLP tasks. Their performance significantly exceeded general BERT despite using fewer labeled examples.

Unsupervised Domain Adaptation

When labels are unavailable in the target domain, techniques from unsupervised domain adaptation can help: adversarial domain alignment (learn representations that are domain-invariant by fooling a domain classifier), self-training (pseudo-label target examples with a source-trained model, then retrain), or test-time adaptation (adapt batch normalization statistics at inference time using unlabeled target samples).

Few-Shot and Zero-Shot Learning

Modern large language models exhibit striking few-shot and zero-shot capabilities — they can perform new tasks from just a handful of examples in the prompt, or even from a task description alone, without any gradient updates. This is fundamentally different from traditional fine-tuning: the model generalizes at inference time, not at training time.

Zero-Shot

A zero-shot approach provides only a task description (and optionally a format instruction) in the prompt, with no examples. GPT-4, Claude, and similar models can answer questions, translate text, and classify sentiment zero-shot because their pretraining data covered these tasks implicitly. Performance depends heavily on the quality of the task description (prompt engineering).

Few-Shot In-Context Learning

Few-shot in-context learning (ICL) provides k labeled examples in the prompt before the test input. The model conditions its generation on these demonstrations, effectively "learning" the task pattern from context. Critically, no gradient updates occur — the weights are frozen. ICL performance scales with model size and is sensitive to example selection and ordering.

Prompt Engineering vs. Fine-Tuning

For large models accessible only via API, prompt engineering (zero/few-shot ICL) may be the only option. For self-hosted models with sufficient data (>100–1000 labeled examples), fine-tuning almost always outperforms ICL on the target task — at the cost of training compute and memory. The choice depends on data availability, compute budget, and latency requirements.

LoRA: Low-Rank Adaptation

Full fine-tuning of large language models (7B–70B parameters) is prohibitively expensive: storing optimizer states for 7B parameters requires ~100 GB of GPU memory. LoRA (Low-Rank Adaptation of Large Language Models) is a parameter-efficient fine-tuning (PEFT) method that dramatically reduces this cost by representing weight updates as low-rank decompositions.

How LoRA Works

For a pretrained weight matrix W₀ ∈ ℝ^{d×k}, LoRA freezes W₀ and adds a trainable low-rank bypass: two small matrices A ∈ ℝ^{d×r} and B ∈ ℝ^{r×k} where r ≪ min(d, k). During the forward pass, the effective weight is W₀ + BA. During training, only A and B are updated; W₀ remains frozen.

LoRA Update
h = W_0 x + \Delta W x = W_0 x + \frac{\alpha}{r} B A x
W₀ is frozen; A is initialized with Gaussian noise, B with zeros (so ΔW = BA = 0 at initialization). α is a scaling hyperparameter; r is the rank (typically 4–64).

The number of trainable parameters is r(d + k) instead of dk. For a typical attention matrix with d = k = 4096 and r = 16, LoRA uses 131K parameters instead of 16.7M — a 128× reduction. LoRA is typically applied to the query and value projection matrices in every attention layer, though it can also be applied to MLP layers.

QLoRA: Quantized LoRA

QLoRA extends LoRA by quantizing the frozen base model weights to 4-bit NF4 precision, reducing memory for a 7B model from ~14 GB (BF16) to ~4 GB. LoRA adapters are trained in BF16 on top of the quantized base. This enables fine-tuning 33B–70B parameter models on a single consumer GPU (e.g., RTX 3090/4090 with 24 GB VRAM) — a dramatic democratization of large model fine-tuning.

Other PEFT Methods

Adapter Layers
Bottleneck modules
Insert small bottleneck modules (down-project → nonlinearity → up-project) between transformer layers. Only adapters are trained. Adds latency at inference.
Prefix Tuning
Trainable prefix tokens
Prepend trainable "prefix" tokens to the key and value sequences in every attention layer. No weight changes; task-specific context is injected via attention.
Prompt Tuning
Soft prompt tokens
Learn a small set of continuous (soft) prompt embeddings prepended to the input. Simpler than prefix tuning; competitive at large scale (11B+).
DoRA
Decomposed LoRA
Decomposes weight into magnitude + direction; fine-tunes both separately. Often outperforms LoRA with similar parameter count.

Practical Fine-Tuning Workflow

A systematic workflow for fine-tuning a pretrained model on a new task:

  1. Choose a base model — pick the smallest model that is likely to have sufficient capacity for the task. Larger is not always better when data is limited.
  2. Inspect and clean your dataset — data quality matters far more than quantity for fine-tuning. Remove duplicates, fix label noise, balance classes if needed.
  3. Decide on strategy — feature extraction for very small datasets; LoRA for large models; full fine-tuning for medium models with sufficient data.
  4. Set learning rate — start with 1e-4 for LoRA adapters, 1e-5 to 5e-5 for full fine-tuning. Use a warmup schedule (5–10% of steps) and cosine decay.
  5. Monitor validation loss — fine-tuning typically converges in 1–5 epochs. Use early stopping to prevent overfitting.
  6. Evaluate on held-out set — use the same metrics as the original task. Check for distributional shift artifacts (e.g., model memorizing dataset artifacts).
  7. Merge and deploy — for LoRA, merge adapters into the base model weights (model.merge_and_unload()) for zero-overhead inference.

When Transfer Learning Fails

Transfer learning is not universally beneficial. It can hurt performance when:

Always run a baseline with feature extraction before committing to expensive full fine-tuning, and compare against a simple in-domain model (e.g., TF-IDF + logistic regression for text) to verify that transfer learning actually helps.


Key Takeaways

Transfer learning reuses representations from large pretrained models to achieve strong performance with little target data. Feature extraction freezes the backbone and trains only a new head — fast and safe for small datasets. Full fine-tuning adapts all weights with a low learning rate — more powerful but risks catastrophic forgetting. LoRA and QLoRA enable parameter-efficient fine-tuning of LLMs on consumer hardware by training low-rank weight updates. Zero-shot and few-shot in-context learning leverage LLMs without any gradient updates. Domain-adaptive pretraining bridges large domain gaps. Choose the simplest strategy that works — not the most complex.

Previous M11-L1: Training at Scale Module Overview Next Lesson M11-L3: Generative Models