A matrix is one of the most powerful and ubiquitous structures in mathematics. At the most basic level, it's just a rectangular array of numbers arranged in rows and columns. But this simple idea encodes something far richer: a matrix represents a linear transformation — a rule for moving, rotating, scaling, or shearing all points in space simultaneously.
A matrix A with m rows and n columns is called an m × n matrix. Its entries are written as a_{ij}, where i is the row index and j is the column index:
A few special names: a square matrix has m = n. A matrix with one column (n = 1) is a column vector. A matrix with one row (m = 1) is a row vector. A 1×1 matrix is a scalar.
The most tangible way to understand matrices is as organized data. Each row might represent an observation; each column a feature. A spreadsheet with 1,000 rows and 50 columns is a 1000×50 matrix.
Example — student grades: Suppose you have 4 students, each with scores in 3 subjects (Math, Physics, English). A 4×3 matrix holds all the data. Row 2 contains student 2's scores. Column 3 contains all students' English scores. Matrix notation gives you a concise language for this structure.
Here's the deeper view: every m×n matrix A defines a function that takes n-dimensional vectors as input and produces m-dimensional vectors as output. This function is linear — it preserves addition and scalar multiplication.
Multiplication of matrix A by vector x gives a new vector Ax. This is the fundamental operation of linear algebra:
The identity matrix has 1s on the diagonal and 0s elsewhere. Multiplying any matrix A by I leaves A unchanged — just like multiplying a number by 1. For any vector x, I·x = x. The identity is the "do nothing" transformation.
The zero matrix has all entries equal to 0. Multiplying any vector by the zero matrix gives the zero vector — everything collapses to the origin. Adding the zero matrix to any matrix A gives A.
A diagonal matrix has nonzero entries only on the main diagonal. Multiplying a diagonal matrix by a vector scales each component independently — it's the simplest non-trivial transformation.
A matrix A is symmetric if A = Aᵀ (it equals its own transpose). Many important matrices in physics and ML are symmetric: covariance matrices, adjacency matrices of undirected graphs, positive definite matrices.
The transpose of matrix A, written Aᵀ, swaps rows and columns: entry (i,j) of A becomes entry (j,i) of Aᵀ. If A is m×n, then Aᵀ is n×m.
Transpose properties: (A+B)ᵀ = Aᵀ + Bᵀ, (cA)ᵀ = cAᵀ, (AB)ᵀ = BᵀAᵀ (note reversed order!), (Aᵀ)ᵀ = A. The reversed-order rule for products is often surprising to newcomers.
Matrices appear in virtually every quantitative field because they efficiently encode relationships between multiple quantities simultaneously: