Reading
Stories Mode

Matrix Operations

~16 min read Lesson 2 of 4 in Module 2

From Structure to Action

In the previous lesson, we learned what a matrix is — a rectangular grid of numbers representing data, transformations, or systems of equations. But a static grid is only half the story. The real power of matrices comes from what we can do with them. Matrix operations turn matrices from passive containers into active engines that drive computation in machine learning, computer graphics, physics simulations, and virtually every domain of quantitative science.

This lesson covers the fundamental operations: addition, scalar multiplication, matrix–vector products, matrix–matrix multiplication, the transpose, and the trace. Each operation has both a mechanical definition and a deeper geometric meaning. Understanding both will make the mechanics feel inevitable rather than arbitrary.

Matrix Addition

The simplest matrix operation is addition. To add two matrices, you simply add their corresponding entries. The entry in row i and column j of A + B is just the sum of the (i,j) entries of A and B. This is called component-wise or entry-wise addition.

There is one constraint: both matrices must have the same dimensions. You cannot add a 3×2 matrix to a 2×3 matrix — the shapes must match exactly. Subtraction works the same way: subtract corresponding entries. The result always has the same dimensions as the inputs.

Matrix Addition
(A+B)_{ij} = a_{ij} + b_{ij}
The (i,j) entry of A+B equals the sum of the (i,j) entries of A and B.
Properties of Matrix Addition

Matrix addition is commutative: A + B = B + A. It is also associative: (A + B) + C = A + (B + C). The zero matrix (all entries zero) acts as an additive identity: A + 0 = A. Every matrix has an additive inverse: A + (−A) = 0.

Scalar Multiplication

To multiply a matrix by a scalar (a single number), you multiply every entry by that scalar. If A is m×n and c is a scalar, then cA is the m×n matrix where each entry (cA)ij = c · aij. This operation has no dimension constraint — it works for any matrix and any scalar.

Geometrically, scalar multiplication scales the transformation represented by the matrix. Multiplying by 2 doubles all the distances; multiplying by −1 reflects through the origin; multiplying by 0 collapses everything to the zero matrix.

Scalar Multiplication
(cA)_{ij} = c \cdot a_{ij}
Every entry is multiplied by the same scalar c — the matrix is uniformly scaled.

Matrix–Vector Product: The Column Combination View

The matrix–vector product Ax is the most important operation in linear algebra. Given an m×n matrix A and an n×1 vector x, the product Ax is an m×1 vector. There are two equivalent ways to compute it, and understanding both gives deep insight.

The row view: each entry of the result is the dot product of the corresponding row of A with the vector x. This is the mechanical computation.

The column view: Ax is a linear combination of the columns of A, with the entries of x as coefficients. This interpretation is geometrically powerful — it reveals that Ax "mixes" the columns of A using the weights in x. Every linear transformation can be understood this way.

Matrix–Vector Product (Column View)
A\mathbf{x} = x_1\mathbf{a}_1 + x_2\mathbf{a}_2 + \cdots + x_n\mathbf{a}_n
Ax is a linear combination of A's columns, weighted by the entries of x.

Matrix–Matrix Multiplication

Matrix multiplication generalizes the matrix–vector product. To multiply A (m×k) by B (k×n), the inner dimensions must match: A has k columns, B has k rows. The result AB is an m×n matrix.

The (i,j) entry of AB is the dot product of row i of A with column j of B. You can also think of it as: the j-th column of AB equals A times the j-th column of B. This means each column of the product is a linear combination of A's columns — the column view again.

Matrix Multiplication Entry Formula
(AB)_{ij} = \sum_{k=1}^{p} a_{ik}\, b_{kj}
Sum over k: row i of A dotted with column j of B gives entry (i,j) of AB.
Dimension Check

Always check dimensions before multiplying. For AB to be defined: the number of columns in A must equal the number of rows in B. If A is m×k and B is k×n, the result is m×n. A mnematch in the inner dimension means the product does not exist.

Order Matters: Non-Commutativity

One of the most important differences between matrix multiplication and ordinary number multiplication is that matrices are not commutative. In general, AB ≠ BA. This is not an exception — it is the rule. Even when both products are defined and have the same dimensions, they will usually be different matrices.

However, matrix multiplication is associative: (AB)C = A(BC). This means you can regroup parentheses without changing the result — but you cannot swap the order of the factors. Multiplication is also distributive: A(B + C) = AB + AC and (A + B)C = AC + BC.

The Transpose

The transpose of a matrix A, written AT, is formed by swapping rows and columns. The entry in row i, column j of A becomes the entry in row j, column i of AT. If A is m×n, then AT is n×m.

The transpose has a beautiful product rule: the transpose of a product reverses the order. This rule generalizes: (ABC)T = CTBTAT. The reversal is necessary — if you forget it, the dimensions won't match.

Transpose Product Rule
(AB)^T = B^T A^T
The transpose of a product reverses the order of the factors.

A matrix is called symmetric if AT = A. Symmetric matrices arise constantly in statistics, physics, and optimization — covariance matrices, the Hessian in calculus, and the Gram matrix XTX are all symmetric. They have special properties that make them easier to work with.

The Trace

The trace of a square matrix is the sum of its diagonal entries. Despite its simplicity, the trace appears in advanced topics like eigenvalue theory, optimization (gradient of a trace), and quantum mechanics.

Trace Definition
\operatorname{tr}(A) = \sum_{i=1}^{n} a_{ii}
The trace is the sum of the diagonal entries — defined only for square matrices.

A key property: trace(AB) = trace(BA), even though AB ≠ BA in general. This cyclic property of the trace is used constantly in matrix calculus and machine learning. The trace also equals the sum of the matrix's eigenvalues — a deep connection that appears in spectral theory.

Computational Cost

Matrix multiplication is the workhorse of modern computing, but it is expensive. Multiplying an m×k matrix by a k×n matrix requires O(mkn) arithmetic operations — roughly m × k × n multiply-add pairs. For square n×n matrices, this is O(n3).

For large matrices, this cost becomes significant. A 1000×1000 matrix multiplication requires about 109 operations. This is why modern deep learning relies on GPUs, which can execute thousands of multiply-adds in parallel. Researchers have also found algorithms that slightly beat O(n3) — the fastest known approaches are around O(n2.37) — though these are rarely used in practice due to large constant factors.

Why These Operations Matter

These operations are not abstract exercises — they are the computational primitives of modern science and engineering:

Neural networks: every layer applies an affine transformation Ax + b, where A is the weight matrix and b is the bias vector. Training updates A and b to minimize a loss function. The entire forward pass through a deep network is a sequence of matrix multiplications.

Computer graphics: 3D rotations, scaling, and perspective projections are all linear transformations represented by matrices. Rendering a 3D scene involves multiplying every vertex by a series of transformation matrices.

Statistics: the sample covariance matrix of a dataset X (with rows as observations) is proportional to XTX. This symmetric matrix encodes how variables vary together — it is central to PCA, linear regression, and many other methods.

Signal processing: convolution, filtering, and Fourier transforms can all be expressed as matrix products. The DFT (Discrete Fourier Transform) is a matrix–vector product with a specific complex matrix.

In the next lesson, we will explore linear transformations in depth — understanding matrices as functions that map vectors to vectors, and building the geometric intuition that ties all of linear algebra together.

Key Takeaways
  • Matrix addition is entry-wise and requires the same dimensions. It is commutative and associative.
  • Scalar multiplication scales every entry by the same number — no dimension constraint.
  • Ax is a linear combination of A's columns weighted by x — the foundation of linear transformations.
  • Matrix multiplication requires matching inner dimensions. Entry (i,j) of AB is the dot product of row i of A and column j of B.
  • Matrix multiplication is NOT commutative (AB ≠ BA in general), but IS associative and distributive.
  • The transpose swaps rows and columns. (AB)T = BTAT — the order reverses.
  • The trace is the sum of diagonal entries. trace(AB) = trace(BA) by the cyclic property.
  • Square n×n matrix multiplication costs O(n3) — GPUs accelerate this via parallelism.
Previous What is a Matrix? Overview Next Linear Transformations