This sandbox was generated by Claude (Anthropic) for the ML 101 course materials. It applies a fixed, named convolution kernel to a small grayscale image and shows the four things the Module 7 lesson describes in words: the feature map a convolution produces, the effect of ReLU, the down-sampling of max-pooling, and how the spatial dimensions shrink at each stage. Nothing here is sketched — every feature map is drawn from the computed array, cell by cell.
These kernels are fixed, not learned. A real CNN learns its kernels by gradient descent over training data; this page does not train anything. It uses the standard textbook operators — Sobel (edges), the Laplacian, a box blur, a sharpen kernel and the identity — so you can see exactly what a convolution does, which is the prerequisite for understanding what a CNN learns. Nothing here is presented as a learned filter.
The convolution of deep learning is cross-correlation. It does not flip the kernel — the flip only changes the result for asymmetric kernels (the Sobel pair) and is a naming convention. This page follows the CNN convention (PyTorch, TensorFlow), and the equation block states it.
Computed: the convolution of the chosen kernel over the chosen image, the ReLU, the 2×2 max-pool, every output dimension, the response energy of each kernel, and the optional second convolution layer — all real arithmetic. Published figures: none — the kernels are standard operators and everything else is a direct computation. Chosen rather than computed: the synthetic grayscale patterns, the image size, and the seed for the optional pixel noise.
Course demo — linked from the Module 7 lesson deck; the page itself is
English‑only for now. A convolution slides a small kernel over an image and, at every
position, takes the weighted sum of the pixels under it. This page lets you pick the kernel, the
padding and the stride, and then computes the whole feature map, its ReLU and its pooling from
the image. One thing to take away: a convolution shrinks the spatial size by
k−1 per valid layer and pooling halves it — which is why deep CNNs
trade spatial resolution for depth.