Networks as Compositions of Linear Maps
A feedforward neural network with L layers maps an input vector x ∈ ℝⁿ to an output ŷ ∈ ℝᵐ by alternating linear transformations and element-wise nonlinearities. Each layer ℓ has a weight matrix W⁽ˡ⁾ ∈ ℝ^{nₗ × nₗ₋₁} and a bias vector b⁽ˡ⁾ ∈ ℝ^{nₗ}. The forward pass at layer ℓ is:
Because each layer applies a matrix multiplication followed by a nonlinearity, the network computes a composition of functions. The power of depth comes from this composition: each layer can represent a different feature detector, and together they build up complex representations from simple ones. Without the nonlinearity σ, any depth-L network would collapse to a single matrix multiplication — depth would give no expressive advantage.
Backpropagation as the Jacobian Chain Rule
Training a neural network means finding weights {W⁽ˡ⁾, b⁽ˡ⁾} that minimize a loss function ℒ(ŷ, y). Gradient descent requires ∂ℒ/∂W⁽ˡ⁾ and ∂ℒ/∂b⁽ˡ⁾ for every layer. Backpropagation computes these gradients efficiently by applying the multivariate chain rule backwards through the network.
The key object is the error signal δ⁽ˡ⁾ = ∂ℒ/∂z⁽ˡ⁾, the gradient of the loss with respect to the pre-activations at layer ℓ. Starting from the output layer and propagating backwards:
In matrix calculus terms, each layer's backward pass multiplies by the Jacobian of that layer's function. For z⁽ˡ⁾ = W⁽ˡ⁾a⁽ˡ⁻¹⁾ + b⁽ˡ⁾, the Jacobian with respect to a⁽ˡ⁻¹⁾ is W⁽ˡ⁾; the chain rule therefore multiplies the upstream gradient by (W⁽ˡ⁾)ᵀ. For a⁽ˡ⁾ = σ(z⁽ˡ⁾), the Jacobian is diag(σ'(z⁽ˡ⁾)) — backpropagation through a nonlinearity is just element-wise multiplication by the derivative.
In the forward pass, W⁽ˡ⁾ maps activations of dimension nₗ₋₁ to pre-activations of dimension nₗ. In the backward pass, we need to map gradients from dimension nₗ back to dimension nₗ₋₁ — which is exactly what (W⁽ˡ⁾)ᵀ does. The transpose reverses the direction of the linear map, propagating error signals upstream through the network architecture.
Computational Graph Perspective
Modern deep learning frameworks (PyTorch, JAX) implement backpropagation via automatic differentiation on a computational graph. Every operation (matrix multiply, activation, loss) is a node; edges carry tensors. Forward pass: compute node values. Backward pass: traverse edges in reverse, accumulating gradients using the chain rule at each node. The network architect only defines the forward computation; gradients are computed automatically.
Batch Normalization
Deep networks are notoriously hard to train: gradients can vanish or explode through many layers, and the distribution of each layer's inputs shifts as earlier layers' parameters change — a problem called internal covariate shift. Batch normalization (Ioffe & Szegedy, 2015) addresses this by normalizing each layer's pre-activations to have zero mean and unit variance within each mini-batch:
From a linear algebra perspective, batch normalization is a linear operation on the activations: subtract the mean vector and divide by the standard deviation vector (both computed over the batch dimension), then apply a learnable diagonal linear transform diag(γ) followed by a translation b = β. The entire transformation is differentiable, so backpropagation works unchanged. During inference, batch statistics are replaced by running estimates accumulated during training.
The Attention Mechanism
Classical neural networks process each input position independently — information can only mix through sequential layers. The attention mechanism (Bahdanau et al., 2015; Vaswani et al., 2017) allows any position to directly attend to any other, learning weighted averages over the input sequence. The core operation uses three matrices: queries Q, keys K, and values V.
Given an input sequence X ∈ ℝ^{n×d} (n tokens, d-dimensional embeddings), the query, key, and value matrices are formed by linear projections: Q = XWᴼ, K = XWᴷ, V = XWⱽ, where Wᴼ, Wᴷ, Wⱽ ∈ ℝ^{d×dₖ} are learned projection matrices. The scaled dot-product attention is then:
The linear algebra is clean: QKᵀ is a matrix of dot products, softmax normalizes each row, and the final multiplication by V computes weighted averages of value vectors. The output at position i is a convex combination of all value vectors, weighted by how much query i matches each key. Positions with high alignment contribute more to the output.
Multi-Head Attention
A single attention head computes one type of relationship (e.g., subject-verb agreement). Multi-head attention runs h independent attention operations in parallel, each with its own projection matrices:
Transformers as Linear Algebra Powerhouses
The transformer architecture (Vaswani et al., 2017) stacks multi-head attention with position-wise feedforward networks, layer normalization, and residual connections. From a linear algebra perspective, every operation is either:
- Matrix multiplication: the QKV projections, the output projection, the feedforward layers (all are linear maps)
- Normalization: layer norm or batch norm (mean centering + diagonal scaling)
- Nonlinearity: softmax in attention, ReLU/GELU in feedforward layers
- Addition: residual connections (x ← x + sublayer(x))
The residual connection x ← x + sublayer(x) is particularly elegant: it means each layer adds a correction to the identity map, rather than computing a new representation from scratch. This makes gradient flow trivial — gradients can flow directly from output to input through the identity branch — and makes training deep transformers (hundreds of layers) practical.
If Q and K have entries drawn from a standard normal distribution, the dot product qᵢᵀkⱼ has variance dₖ (sum of dₖ products of unit-variance variables). For large dₖ, this pushes softmax into regions where one score dominates and gradients vanish. Dividing by √dₖ restores unit variance, keeping attention weights spread over multiple keys and gradients healthy throughout training.
Putting It Together: A Layer as a Linear Map
Every component of a transformer can be understood as a learned linear map (or a composition of linear maps and pointwise operations). The weight matrices are the parameters — learned by gradient descent to minimize a task loss. Backpropagation computes gradients of the loss with respect to every weight matrix by applying the chain rule layer by layer, propagating error signals backward through transposed weight matrices.
The remarkable fact is that all of this — forward pass through billions of parameters, backward pass computing billions of gradients — reduces to batched matrix multiplications, the operation that modern GPUs/TPUs execute most efficiently. Linear algebra is not incidental to deep learning; it is the substrate on which deep learning runs.
A neural network is a composition of matrix multiplications and element-wise nonlinearities. Forward propagation applies these maps in sequence: z⁽ˡ⁾ = W⁽ˡ⁾a⁽ˡ⁻¹⁾ + b⁽ˡ⁾, a⁽ˡ⁾ = σ(z⁽ˡ⁾). Backpropagation computes gradients via the chain rule, propagating error signals δ⁽ˡ⁾ backwards through the transposed weight matrices. Batch normalization normalizes pre-activations to zero mean and unit variance within each mini-batch, using learnable scale γ and shift β. The attention mechanism computes scaled dot-product similarities between queries and keys, producing weighted sums of values: Attention(Q,K,V) = softmax(QKᵀ/√dₖ)V. Multi-head attention runs h independent attention operations in parallel, each learning different relationship types. Transformers stack these components with residual connections — enabling gradient flow through hundreds of layers — making the entire architecture expressible as sequences of matrix operations.