Neural Network Fundamentals
Neural networks are chains of matrix multiplications. Forward propagation composes linear maps; backpropagation differentiates them via the Jacobian chain rule. Attention weighs relevance with dot-product similarity.
Layer by Layer
Each layer applies a weight matrix W⁽ˡ⁾ and bias b⁽ˡ⁾, then a nonlinearity σ. The full forward pass is a composition of L affine maps separated by nonlinearities.
Composition, Not Width
Without nonlinearities, any L-layer network collapses to a single matrix multiply — depth adds no expressive power. Nonlinearities break linearity, letting each layer detect increasingly abstract features: edges → textures → objects.
Chain Rule with Jacobians
Error signal δ⁽ˡ⁾ = ∂ℒ/∂z⁽ˡ⁾ propagates backwards through transposed weight matrices. The weight gradient is the outer product δ⁽ˡ⁾(a⁽ˡ⁻¹⁾)ᵀ.
Reversing the Direction
Forward pass: W maps nₗ₋₁ → nₗ. Backward pass: gradients must flow nₗ → nₗ₋₁. (W)ᵀ reverses the map — that is exactly what we need to propagate error signals back through the architecture.
Centering the Activations
Normalize pre-activations to zero mean, unit variance within each mini-batch. Add learnable scale γ and shift β so the network can undo normalization if needed. Makes gradient flow stable through deep networks.
Queries, Keys, Values
Project input X into Q, K, V via learned matrices. Compute scaled dot-product attention: each position queries all positions, weights them by similarity, and averages the values.
Variance Control
Dot products qᵢᵀkⱼ have variance dₖ. Large dₖ → saturated softmax → vanishing gradients. Dividing by √dₖ restores unit variance, keeping attention weights spread and gradients healthy.
Parallel Attention Heads
Run h attention operations in parallel, each with its own learned projections. Concatenate outputs and project. Different heads learn different relationship types simultaneously: syntax, coreference, semantics.
- Head 1: syntactic dependency relations
- Head 2: coreference resolution
- Head h: positional patterns
Key Takeaways
- Forward pass: z⁽ˡ⁾ = W⁽ˡ⁾a⁽ˡ⁻¹⁾ + b⁽ˡ⁾, a⁽ˡ⁾ = σ(z⁽ˡ⁾) — compositions of linear maps
- Backprop: δ⁽ˡ⁾ = (W⁽ˡ⁺¹⁾)ᵀδ⁽ˡ⁺¹⁾ ⊙ σ'(z⁽ˡ⁾) — Jacobian chain rule backwards
- Batch norm: normalize to μ=0, σ²=1 per mini-batch; learnable γ, β
- Attention: softmax(QKᵀ/√dₖ)V — weighted average of values by query-key similarity
- Transformers: multi-head attention + residual connections = practical deep stacks
Linear Algebra Powers AI
You've seen linear algebra at the heart of machine learning: regression, matrix factorization, and now neural networks. Every computation — forward pass, backprop, attention — is matrix operations at scale.