This sandbox was generated by Claude (Anthropic) for the ML 101 course materials. It
demonstrates the transfer-learning mechanism on a small 2D task with one model class —
x → tanh(W₁x + b₁) → w·h + b₀ — in three regimes that differ only
in how they start and what is frozen: from scratch (random init, train everything),
frozen features (fix the hidden layer, train only a linear head), and fine-tune
(start from the frozen values, then train everything). It shows the result the Module 11
lesson describes in words: with few labeled samples the frozen extractor plus a small head wins,
and the advantage fades as labels grow.
There is no real pretrained model here, and the page does not pretend there is. The fixed hidden layer is a seeded random-feature map — tanh random features approximating an RBF kernel, with a bandwidth chosen to suit this task family — and it is the one illustrative stand-in for “features a big model would have learned on a large related dataset”. The bandwidth is the only “pretraining knowledge” it carries, and you can turn it.
Everything else is computed, live and seeded. The head is a closed-form ridge solve on the frozen features; the from-scratch and fine-tuned networks are trained by real full-batch gradient descent with backpropagation; every accuracy is measured on a fixed held-out test set; and the sample-efficiency curve is averaged over several seeded label draws. Change the seed and the exact numbers move; the ordering at small n does not.
Computed: the feature map (seeded, reproducible), the ridge head, the two gradient-descent networks, every train and test accuracy, the accuracy-vs-n curve, the sample-efficiency gap, and the decision boundaries by evaluating each model over a grid. Published figures: none. Chosen rather than computed: the four task shapes, the feature dimension and bandwidth defaults, the noise level, the seeds, the training budget, and the target accuracy used to read “labels needed”.
Course demo — linked from the Module 11 “Transfer Learning & Fine-Tuning” lesson; the page itself is English‑only for now. Transfer learning reuses a feature extractor learned on a big related dataset and trains only a small head on your task. This page stands that extractor in with a fixed seeded feature map and then computes how a frozen head, a fine-tuned network and a from-scratch network compare as you change the number of labeled samples. One thing to take away: with few labels the frozen good features win, and the gap shrinks as labels grow — the whole reason few-shot transfer works.