Labs
Demos
Prevalence, cost and the threshold
Hold the model fixed and change only how rare the positives are. AUC does not move; precision, F1 and PR-AUC collapse. Also prices false positives against false negatives to find the threshold that minimises expected cost.
Bias-variance sandbox
Raise the model complexity and watch training error keep falling while test error turns upward. Bias and variance are COMPUTED over many training sets, not described: the fan of fitted curves is what variance actually looks like.
Decision-tree sandbox
Build a decision tree on a small 2D dataset: Gini impurity, entropy and the information gain of every candidate split are computed, and the axis-aligned decision boundary is drawn by running the tree's own predict() over a grid. Set the depth, min-samples and criterion, and watch it overfit noise.
SVM visualizer
Train a soft-margin SVM with a real dual solver (simplified SMO): the weight vector, bias, margin width 2/||w|| and support vectors are all solved for, and the boundary and margins are drawn from the decision function. Turn C, switch to the RBF kernel and tune gamma, and watch the margin, the support vectors and the C-sweep respond.
Clustering explorer
Run a real K-means (Lloyd with a k-means++ start) or DBSCAN on a seeded 2D dataset: the assignment, centroids, core/border/noise labels, inertia and silhouette are all computed. Switch to the moons or rings to watch K-means cut across a shape DBSCAN wraps, and read the elbow and the k-distance graph.
CNN feature visualizer
Apply a fixed, named convolution kernel (Sobel, Laplacian, blur, sharpen, identity) to a small grayscale image, then ReLU, then 2x2 max-pool: every feature map is real 2D-convolution arithmetic, never sketched. Set padding, stride and a second conv layer, and watch the spatial dimensions shrink stage by stage while each kernel detects edges, smooths or sharpens.
Gradient descent on a live loss surface
Watch descent diverge, oscillate or crawl on a live loss surface. The learning-rate stability law is derived from the Hessian and stated on screen.
Backpropagation step-through
One backpropagation step at a time on a tiny network, with every partial derivative shown as it is computed.
Genetic algorithm, one generation at a time
One generation at a time: selection, crossover and mutation applied to a real population, with the fitness distribution redrawn each step.
Q-learning, one update at a time
One Q-update at a time on a grid world, with the Bellman arithmetic shown for the state the agent is actually in.
Sequence-prediction sandbox
Build a character-level n-gram language model on a built-in text corpus: the add-k smoothed next-character distribution, the training and held-out perplexity (in bits), and seeded sample text are all computed from real counts -- no neural network. Set the order, smoothing k and temperature, drop k to 0 and watch held-out perplexity blow up, and read the generated text as texture.
Transfer Learning Workshop
Demonstrates the transfer-learning MECHANISM on a small 2D task: a fixed seeded feature map (a synthetic STAND-IN for features a big model would have learned on a large related dataset, NOT real pretrained weights) is frozen and only a linear head is trained, versus a network trained from scratch and one fine-tuned from the frozen start. The head, both networks, every accuracy and the accuracy-vs-labels curve are computed; with few labels the frozen features win and the gap shrinks as labels grow.