The Word We Have Been Using Loosely
Three lessons in, you have read the word “space” a dozen times — vector space, 2D space, 3D space — and nobody has said what it means. That is about to become a problem. From Module 3 onward, almost every result in this course is phrased in terms of subspaces: the null space of a matrix, the column space, an eigenspace, the space a projection lands in. If “subspace” is vague for you, every one of those is vague too.
So this lesson does the unglamorous work. It defines a vector space, defines a subspace, hands you a three-line test you can actually run, and connects all of it back to the pictures from Lesson 3. Then it closes a gap: a norm you have not met yet, the L∞ norm, because “how big is this vector” turns out to have more than one right answer.
What a Vector Space Is
A vector space is a set V, together with two operations, satisfying a short list of rules. The two operations are ones you already use: add two elements of V, and multiply an element of V by a real number (a scalar). The crucial requirement is that both operations stay inside V. Add two elements of V and the result is in V. Scale an element of V and the result is in V. That property is called closure, and it is the beating heart of the definition.
The working example throughout this course is ℝⁿ: all lists of n real numbers. Add two lists componentwise and you get a list of n real numbers. Multiply a list by 2.5 and you get a list of n real numbers. Nothing you can do with addition and scaling will ever produce something outside ℝⁿ. So ℝⁿ is a vector space, and its elements are what we have been calling vectors.
In this framework, vector is not a claim about shape. It means “element of a vector space”. An arrow in ℝ³ is a vector. So is a polynomial, a function, a matrix, or a digitised audio clip — provided it lives in a set where adding and scaling behave properly.
The Axioms, Stated Plainly
The standard list has eight axioms, and every one is something you would have assumed anyway. Read them once, precisely so that you can stop thinking about them.
- Closure under addition — if u and v are in V, so is u + v.
- Closure under scaling — if u is in V and c is real, cu is in V.
- Commutativity — u + v = v + u.
- Associativity — (u + v) + w = u + (v + w).
- Additive identity — there is a zero vector 0 in V with v + 0 = v for every v.
- Additive inverses — every v in V has a −v in V with v + (−v) = 0.
- Distributivity — c(u + v) = cu + cv, and (c + d)u = cu + du.
- Scalar compatibility — c(du) = (cd)u, and 1u = u.
Two consequences fall straight out and are worth noticing. The zero vector is unique, and 0 v = 0 for every v — scaling anything by zero lands you at the origin. That second fact is why the origin shows up in every subspace, which is where we are heading.
A Vector Space With No Arrows In It
If we only ever used ℝⁿ, the axioms would look like ceremony. Here is the example that shows why they are worth stating. Let P₂ be the set of all polynomials of degree 2 or less — every expression of the form a + bx + cx², with a, b and c real.
Add two of them: (1 + 2x) + (3 − x + 4x²) = 4 + x + 4x². Still degree 2 or less, still in P₂. Scale one by −3: −3(1 + 2x) = −3 − 6x. Still in P₂. The zero polynomial plays the role of the zero vector, and every remaining axiom is just ordinary algebra with polynomials. P₂ is a vector space, and there is not an arrow in sight.
Notice something else. P₂ has an obvious minimal spanning set — 1, x and x² — of size three, so it is three-dimensional in exactly the sense ℝ³ is. That correspondence is not a coincidence, and it is why the same theorems serve signal processing, statistics and machine learning. Signal spaces, function spaces and image spaces are all vector spaces, and every result in this course applies to them unchanged.
Subspaces: A Space Inside a Space
Now the definition the rest of the course depends on. Let V be a vector space. A subset S of V is a subspace of V if S is itself a vector space, using the same addition and the same scalar multiplication that V already uses.
That last clause is doing real work. We are not building a new structure; we are asking whether a piece of an existing one is self-contained. And self-contained means exactly closure: take elements of S, add them, scale them, and you must never land outside S. A subspace is a region you cannot escape by doing linear algebra.
There are exactly four kinds: the single point {0}; any line through the origin; any plane through the origin; and all of ℝ³. That is the complete list. A sphere is not a subspace. A line that misses the origin is not a subspace. The first octant is not a subspace. Looking like a tidy set is not the criterion.
The Three-Part Test
Checking eight axioms for every candidate subset would be miserable, and it is unnecessary. Commutativity, associativity and distributivity are inherited for free — they hold for elements of S because those elements are also elements of V, where the axioms already hold. Only closure can fail. So the whole question reduces to three checks.
In words: the zero vector of V is in S; if u and v are in S then u + v is in S; and if u is in S and c is any real number then cu is in S. All three hold and S is a subspace. Any one fails and S is not a subspace — there is no partial credit.
Check the zero vector first. It is the cheapest of the three and the one that fails most often, which makes it an efficient way to be finished early. Strictly, condition three with c = 0 already forces 0 into S whenever S has anything in it at all, so condition one is really doing double duty as a non-emptiness check — but as a practical first move it is the best you have.
Example 1: The xy-Plane Inside ℝ³
Let V = ℝ³ and let S be every vector whose third coordinate is zero. Geometrically that is the xy-plane, sitting inside 3D space and passing through the origin. Run the test.
Zero? (0, 0, 0) has third coordinate 0, so yes. Sums? (x₁, y₁, 0) + (x₂, y₂, 0) = (x₁ + x₂, y₁ + y₂, 0) — the third coordinate is 0 + 0 = 0, so the sum is still in S. Scalings? c(x, y, 0) = (cx, cy, 0), and c · 0 = 0, so the result is still in S. All three pass: S is a subspace of ℝ³, and its dimension is 2.
Example 2: A Line Through the Origin
Let V = ℝ² and let S be every real multiple of the vector (2, 1) — the line through the origin with slope ½. Testing this one is almost mechanical, which is exactly the point.
Zero? Take c = 0 and you get (0, 0). Yes. Sums? c₁(2, 1) + c₂(2, 1) = (c₁ + c₂)(2, 1), which is another multiple of (2, 1). Yes. Scalings? k · c(2, 1) = (kc)(2, 1). Yes. So S is a subspace, of dimension 1. Notice what did the work: S was defined as “all multiples of a vector”, and closure came for free. Hold onto that — it returns in a few paragraphs as the span.
The Non-Example: A Line That Misses the Origin
Now one that fails, and it matters that you watch it fail, because it looks so respectable. In ℝ², let S be the line y = x + 1. It is a perfectly good line: straight, infinite in both directions, no harder to draw than the line in Example 2. It is not a subspace.
Zero? Is (0, 0) on the line? That needs 0 = 0 + 1, which is false. The test is already over. But it is instructive to watch closure fail too. Take (0, 1) and (1, 2), both genuinely on the line. Their sum is (1, 3), and the line would require 3 = 1 + 1 = 2. It does not. Adding two elements of S produced something outside S.
A line or plane is a subspace only if it passes through the origin. A line or plane shifted away from the origin is called an affine set. Affine sets are genuinely useful — the solution set of Ax = b for b ≠ 0 is one — but they are not subspaces, and subspace theorems do not apply to them.
Span: Everything You Can Reach
You met span geometrically in Lesson 3: two non-parallel vectors in 2D span the whole plane, while parallel ones span only a line. Here is the same idea as a definition rather than a picture. The span of a set of vectors is the set of all linear combinations of them — everything reachable by scaling and adding.
And now a theorem you get for free: a span is always a subspace. You never have to test one. Zero is in it (take every coefficient to be zero). A sum of two linear combinations is a linear combination. A scalar times a linear combination is a linear combination. All three checks pass by construction, for any set of vectors in any vector space.
That is why Example 2 was so easy — the line along (2, 1) was already a span. Read the other way, this gives you a way to build subspaces instead of testing them: pick some vectors, take their span, and you have a subspace, guaranteed.
Basis and Dimension
A span can be wasteful. The span of (1, 0), (0, 1) and (1, 1) is all of ℝ², but the third vector adds nothing — it was already reachable from the first two. A basis is a spanning set with no waste: it spans the space, and no vector in it is redundant. Equivalently, remove any vector from a basis and the span shrinks.
Here is the fact that makes the next word meaningful: every basis of the same vector space has the same number of vectors. You can choose different bases — infinitely many of them — and the count never changes. That invariant count is the dimension of the space.
So ℝ³ has dimension 3, a plane through the origin inside it has dimension 2, a line through the origin has dimension 1, and {0} has dimension 0 — the empty set spans it. P₂ from earlier has dimension 3, with basis 1, x, x². Dimension is how you compare spaces that look nothing like each other.
Where Subspaces Pay Off
All of this becomes concrete in Module 3. Given a matrix A, ask which vectors it sends to zero — the solutions of Ax = 0. That set is the null space (or kernel) of A, and the three-part test on it is a two-line exercise. A0 = 0, so the zero vector is in it. If Ax = 0 and Ay = 0 then A(x + y) = 0. And A(cx) = cAx = 0. It is a subspace, and its dimension has a name: the nullity of A.
Meanwhile the columns of A span a subspace of their own — the column space — and its dimension is the rank. The rank–nullity theorem, which ties those two numbers to the width of A, is a statement about the dimensions of two subspaces and nothing more. Module 3, Lesson 3 is where it arrives, and it is the first place the definitions in this lesson do real work.
Measuring Size: L1, L2 and L∞
One loose end before Module 1 closes. Lesson 3 defined the length of a vector as the square root of the sum of its squared components, and Module 11 uses that length constantly under the name L2 norm, alongside a second one called the L1 norm. There is a third member of the same family that gets used just as often and introduced far less: the L∞ norm, also called the max norm.
A norm is any function that assigns a non-negative length to a vector, returns zero only for the zero vector, scales the way you expect, and obeys the triangle inequality. Several different functions satisfy all of that at once, and each measures “how big” in its own way.
Take v = (3, −4, 1). Its L1 norm is 3 + 4 + 1 = 8. Its L2 norm is the square root of 9 + 16 + 1, which is about 5.10. Its L∞ norm is max(3, 4, 1) = 4 — it reports the single worst component and discards the rest. Those three numbers describe the same vector; they answer three different questions about it.
L∞ is what you want whenever the worst case is the thing that matters. “No sample in this block exceeds 4” is a statement about the max norm, not the Euclidean one. It also arrives as a limit: the general Lp norm raises each component to the power p, sums, and takes the p-th root. At p = 1 that is L1; at p = 2 it is L2; and as p grows the largest component dominates the sum completely, so the whole expression converges to max|vᵢ|. That limit is why the notation says infinity.
The Unit Ball: Three Shapes
The clearest way to feel the difference is to draw the unit ball of each norm: every vector whose length is at most 1. In ℝ² all three are convex and symmetric about the origin, and they look nothing alike.
The L2 unit ball is the familiar circle of radius 1. The L1 unit ball is the diamond with corners at (±1, 0) and (0, ±1): starting from (1, 0), you can trade horizontal for vertical one unit at a time, so the boundary comes out straight. The L∞ unit ball is the square with corners at (±1, ±1), because both coordinates are free to reach 1 independently of each other.
Those shapes explain later results without any algebra. The L1 diamond has its corners on the axes, which is why an L1 penalty drives coefficients exactly to zero — the optimum tends to land on a corner, and a corner has a zero coordinate. The L2 circle is smooth everywhere, has no preferred direction, and so produces no exact zeros. The story slides for this lesson let you slide p upward and watch the diamond swell into the circle and then harden into the square.
One inequality falls straight out of the picture: the diamond sits inside the circle, which sits inside the square. So the max norm is never larger than the Euclidean norm, which is never larger than the taxicab norm — for every vector, in every dimension.
- A vector space is a set closed under addition and scalar multiplication, plus six more axioms you already obey without thinking.
- Polynomials, functions and signals form vector spaces too — “vector” is a role, not a shape.
- A subspace is a subset that is a vector space in its own right: a region you cannot escape by adding and scaling.
- The three-part test is the whole job — contains the zero vector, closed under addition, closed under scalar multiplication.
- A line or plane is a subspace only if it passes through the origin; shifted ones are affine, not subspaces.
- Span is everything reachable, and a span is automatically a subspace — which is how you build subspaces rather than test them.
- A basis is a spanning set with nothing spare; every basis of a space has the same size, and that size is the dimension.
- The null space and column space of a matrix are subspaces; their dimensions are the nullity and the rank (Module 3, Lesson 3).
- L1 sums absolute values, L2 is straight-line length, and L∞ takes the largest component — diamond, circle and square unit balls respectively.