LA 101
M11 · L03
Module 11: Machine Learning Applications

Neural Network Fundamentals

Neural networks are chains of matrix multiplications. Forward propagation composes linear maps; backpropagation differentiates them via the Jacobian chain rule. Attention weighs relevance with dot-product similarity.

01 / 12
LA 101
M11 · L03
Forward Propagation

Layer by Layer

Each layer applies a weight matrix W⁽ˡ⁾ and bias b⁽ˡ⁾, then a nonlinearity σ. The full forward pass is a composition of L affine maps separated by nonlinearities.

Forward Pass
z^{(\ell)}=W^{(\ell)}a^{(\ell-1)}+b^{(\ell)},\quad a^{(\ell)}=\sigma(z^{(\ell)})
02 / 12
LA 101
M11 · L03
Why Depth Works

Composition, Not Width

Without nonlinearities, any L-layer network collapses to a single matrix multiply — depth adds no expressive power. Nonlinearities break linearity, letting each layer detect increasingly abstract features: edges → textures → objects.

σ
Nonlinearity
W
Linear map
L
Depth
03 / 12
LA 101
M11 · L03
Backpropagation

Chain Rule with Jacobians

Error signal δ⁽ˡ⁾ = ∂ℒ/∂z⁽ˡ⁾ propagates backwards through transposed weight matrices. The weight gradient is the outer product δ⁽ˡ⁾(a⁽ˡ⁻¹⁾)ᵀ.

Backprop Recurrence
\delta^{(\ell)}=\left(W^{(\ell+1)}\right)^T\delta^{(\ell+1)}\odot\sigma'(z^{(\ell)})
04 / 12
LA 101
M11 · L03
Why the Transpose?

Reversing the Direction

Forward pass: W maps nₗ₋₁ → nₗ. Backward pass: gradients must flow nₗ → nₗ₋₁. (W)ᵀ reverses the map — that is exactly what we need to propagate error signals back through the architecture.

Gradient shapes
∂ℒ/∂W⁽ˡ⁾ = δ⁽ˡ⁾(a⁽ˡ⁻¹⁾)ᵀ · same shape as W⁽ˡ⁾
05 / 12
LA 101
M11 · L03
Batch Normalization

Centering the Activations

Normalize pre-activations to zero mean, unit variance within each mini-batch. Add learnable scale γ and shift β so the network can undo normalization if needed. Makes gradient flow stable through deep networks.

Batch Norm
\hat{x}_i=\frac{x_i-\mu_B}{\sqrt{\sigma_B^2+\varepsilon}},\quad y_i=\gamma\hat{x}_i+\beta
06 / 12
LA 101
M11 · L03
Attention Mechanism

Queries, Keys, Values

Project input X into Q, K, V via learned matrices. Compute scaled dot-product attention: each position queries all positions, weights them by similarity, and averages the values.

Scaled Dot-Product
\text{Attention}(Q,K,V)=\text{softmax}\!\left(\frac{QK^T}{\sqrt{d_k}}\right)V
07 / 12
LA 101
M11 · L03
Why √dₖ?

Variance Control

Dot products qᵢᵀkⱼ have variance dₖ. Large dₖ → saturated softmax → vanishing gradients. Dividing by √dₖ restores unit variance, keeping attention weights spread and gradients healthy.

Var
Without √dₖ: dₖ
1
With √dₖ: 1
08 / 12
LA 101
M11 · L03
Multi-Head Attention

Parallel Attention Heads

Run h attention operations in parallel, each with its own learned projections. Concatenate outputs and project. Different heads learn different relationship types simultaneously: syntax, coreference, semantics.

  • Head 1: syntactic dependency relations
  • Head 2: coreference resolution
  • Head h: positional patterns
09 / 12
LA 101
Neural network fundamentals

Check what stuck

Four questions on forward and backward passes and attention as matrix operations.

Question 1 of 0
Score 0/0

10 / 12
LA 101
Key Takeaways
Summary

Key Takeaways

  • Forward pass: z⁽ˡ⁾ = W⁽ˡ⁾a⁽ˡ⁻¹⁾ + b⁽ˡ⁾, a⁽ˡ⁾ = σ(z⁽ˡ⁾) — compositions of linear maps
  • Backprop: δ⁽ˡ⁾ = (W⁽ˡ⁺¹⁾)ᵀδ⁽ˡ⁺¹⁾ ⊙ σ'(z⁽ˡ⁾) — Jacobian chain rule backwards
  • Batch norm: normalize to μ=0, σ²=1 per mini-batch; learnable γ, β
  • Attention: softmax(QKᵀ/√dₖ)V — weighted average of values by query-key similarity
  • Transformers: multi-head attention + residual connections = practical deep stacks
11 / 12
LA 101
Module Complete
Module 11 Complete

Linear Algebra Powers AI

You've seen linear algebra at the heart of machine learning: regression, matrix factorization, and now neural networks. Every computation — forward pass, backprop, attention — is matrix operations at scale.

12 / 12