Reading
Stories Mode

Random Variables

~16 min read Lesson 1 of 4 in Module 4

Bridging Probability and Measurement

Probability theory gives us tools to reason about uncertainty. But real-world data is numerical — coin flips become 0s and 1s, dice rolls become integers from 1 to 6, heights become measurements in centimeters. The bridge between abstract probability and measurable numbers is the random variable.

A random variable is not actually a variable in the algebraic sense, and it is not random in the colloquial sense of being arbitrary. It is a function that maps each outcome of a random experiment to a real number. This seemingly simple idea unlocks a vast toolkit: expected values, variances, distribution functions, and eventually the entire machinery of statistical inference.

Formal Definition

A random variable X is a function X : Ω → ℝ that assigns a real number X(ω) to each outcome ω in the sample space Ω. The randomness comes from the underlying experiment, not from X itself.

X(ω) = numerical value assigned to outcome ω

Discrete vs. Continuous Random Variables

The most fundamental distinction among random variables is whether they take on a countable set of values or an uncountable continuum of values.

A discrete random variable takes on values from a finite or countably infinite set: the number of heads in ten coin flips, the count of defects on a circuit board, the number of customers who arrive in an hour. You can list the possible values even if the list is infinite.

A continuous random variable takes on values from an uncountable interval: the exact voltage of a signal, the waiting time until the next earthquake, the temperature at noon tomorrow. Between any two possible values there are infinitely many other possible values, so you cannot list them.

Key Distinction

Discrete: counts things. Values are isolated points. Probabilities are assigned to individual values.

Continuous: measures things. Values fill an interval. The probability of any single exact value is zero — only intervals have nonzero probability.

Probability Mass Function (PMF)

For a discrete random variable X, the Probability Mass Function (PMF) tells you the probability that X equals each specific value x. It answers: how much probability mass sits at each point?

Probability Mass Function
p(x) = P(X = x)
p(x) gives the probability that the random variable X takes the specific value x. The PMF must be non-negative and must sum to 1 over all possible values.

The PMF must satisfy two conditions: p(x) ≥ 0 for all x, and the sum of p(x) over all possible values equals 1. These are just the probability axioms translated to the language of functions.

A simple example: roll a fair six-sided die. The random variable X = outcome takes values {1, 2, 3, 4, 5, 6}, each with PMF p(x) = 1/6. The PMF is a flat function — a uniform distribution over six points. More complex random variables, such as the sum of two dice, have non-uniform PMFs with a characteristic tent shape.

Probability Density Function (PDF)

For a continuous random variable X, we cannot assign positive probability to individual points — there are uncountably many of them, and probabilities must sum to 1. Instead, we use a Probability Density Function (PDF), denoted f(x), which describes how probability is spread across the real line.

The PDF f(x) is not a probability itself. Rather, probability is area under the density curve. To find the probability that X falls in the interval [a, b], you integrate the PDF over that interval.

Probability Density Function
P(a \le X \le b) = \int_a^b f(x)\,dx
The probability that X falls in [a, b] is the area under f(x) between a and b. The total area under the PDF must equal 1.

The PDF must satisfy f(x) ≥ 0 everywhere, and the integral over all of ℝ must equal 1. Unlike a PMF, the PDF value f(x) can exceed 1 at a point — it is a density, not a probability. Think of it like mass density: a small region of high density has a lot of probability packed into it.

Cumulative Distribution Function (CDF)

Both discrete and continuous random variables have a Cumulative Distribution Function (CDF), denoted F(x). The CDF answers: what is the probability that X is at most x? It accumulates probability from negative infinity up to x.

Cumulative Distribution Function
F(x) = P(X \le x)
F(x) is non-decreasing, approaches 0 as x → −∞, and approaches 1 as x → +∞. For continuous RVs, F(x) = ∫ f(t) dt from −∞ to x.

The CDF is always non-decreasing because accumulating more probability never decreases the total. It starts at 0 (no probability to the left of everything) and ends at 1 (all probability to the left of everything). For a discrete variable, the CDF is a staircase function that jumps at each possible value. For a continuous variable, it is a smooth S-shaped curve.

The CDF is particularly useful for computing probabilities of intervals. P(a < X ≤ b) = F(b) − F(a). Flipping between the PDF and CDF via differentiation and integration is a fundamental operation in probability theory.

Expected Value

The expected value E[X], also called the mean or expectation, is the long-run average value of a random variable over many independent repetitions of the experiment. It is a weighted average of all possible values, weighted by their probabilities.

Expected Value (Discrete)
E[X] = \sum_x x\,p(x)
Sum of each value times its probability mass
Expected Value (Continuous)
E[X] = \int_{-\infty}^{\infty} x\,f(x)\,dx
Integral of x times the probability density

The expected value need not be a value that X can actually take. If you roll a fair die, E[X] = 3.5, but the die never lands on 3.5. Think of E[X] as the balancing point of the probability distribution — if you placed weights equal to the probabilities on a number line, the expected value is where the line would balance.

The expected value is linear: E[aX + b] = a E[X] + b, and E[X + Y] = E[X] + E[Y] for any random variables X and Y. This linearity is one of the most powerful tools in probability, holding even when X and Y are dependent.

Variance and Standard Deviation

The expected value tells us where the distribution is centered, but not how spread out it is. Two distributions can have the same mean but completely different shapes. The variance Var(X) quantifies this spread by measuring the average squared deviation from the mean.

Variance
\text{Var}(X) = E\!\left[(X-\mu)^2\right] = E[X^2] - (E[X])^2
The computational form E[X²] − (E[X])² is often easier to calculate. The units of variance are the square of the units of X.

Why squared deviations? Using absolute deviations |X − μ| is mathematically less convenient because the absolute value function is not differentiable at zero. Squared deviations lead to formulas that are much more tractable analytically and computationally.

The standard deviation σ = √Var(X) restores the original units of measurement, making it more interpretable. If X represents heights in centimeters, Var(X) is in cm², but σ is in cm — the same unit as the data.

Variance Properties

Var(X) ≥ 0 always. Var(X) = 0 if and only if X is a constant (no spread at all).

Var(aX + b) = a² Var(X) — shifting by b does not change spread, scaling by a squares the variance.

For independent random variables: Var(X + Y) = Var(X) + Var(Y).

Functions of Random Variables

If X is a random variable and g is a function, then Y = g(X) is also a random variable. Its expected value is computed using the law of the unconscious statistician (LOTUS): you do not need to find the distribution of Y first. Instead, you apply g to each value of X and weight by the probability of X.

LOTUS (Discrete)
E[g(X)] = \sum_x g(x)\,p(x)
To find E[g(X)], sum g(x)·p(x) over all values x — no need to derive the distribution of g(X) first

LOTUS is what makes the variance formula work: Var(X) = E[(X − μ)²] is just E[g(X)] where g(x) = (x − μ)². It shows that variance is a special case of expected value applied to a transformed random variable.

Why Random Variables Are the Central Language of Statistics

The random variable abstraction is powerful because it separates the underlying probability model from the specific numbers we observe. Once you have expressed a situation as a random variable, you can apply a common toolkit regardless of what the original experiment was.

Every statistical model is ultimately a statement about random variables. A linear regression model says that the response Y is a linear function of predictors X1, …, Xp plus a random variable ε representing noise. A hypothesis test asks whether the observed test statistic, treated as a random variable under the null hypothesis, is consistent with what we would expect.

In Module 4, we will explore the most important families of random variables — binomial, Poisson, normal, exponential — that appear again and again in practice. Each family is characterized by its PMF or PDF, its expected value, and its variance, all of which follow directly from the definitions we established in this lesson.

Key Takeaways
  • A random variable is a function that assigns a real number to each outcome of a random experiment — it is the bridge between sample spaces and arithmetic.
  • Discrete RVs take countable values; their distribution is described by the PMF p(x) = P(X = x).
  • Continuous RVs take values on a continuum; probability is area under the PDF f(x), not f(x) itself.
  • The CDF F(x) = P(X ≤ x) applies to both types, is always non-decreasing, and runs from 0 to 1.
  • The expected value E[X] is the probability-weighted average of all possible values — the balancing point of the distribution.
  • Linearity of expectation: E[aX + b] = aE[X] + b, and E[X + Y] = E[X] + E[Y] for any X and Y.
  • Variance Var(X) = E[(X − μ)²] measures average squared spread around the mean; standard deviation σ = √Var(X) restores original units.
Previous Bayes’ Theorem in Practice Module Overview Next Common Discrete Distributions