From Counting to Measuring
In the previous lesson we studied distributions for counting things — successes in trials, arrivals in an interval, trials until a success. Now we shift to continuous random variables, which model quantities we measure rather than count: height, time, temperature, price, distance.
For a continuous random variable X, we no longer assign probability to individual values — the probability that X equals exactly 3.14159 is zero. Instead, probability is assigned to intervals via the probability density function (PDF). The PDF f(x) satisfies f(x) ≥ 0 and ∫ f(x) dx = 1, and P(a ≤ X ≤ b) = ∫ab f(x) dx.
In this lesson we survey seven fundamental continuous distributions: Uniform, Normal (Gaussian), Exponential, Gamma, Beta, and Log-normal. Together they underpin an enormous range of engineering, scientific, and data analysis problems.
The Uniform Distribution
The simplest continuous distribution assigns equal probability density to every point in an interval [a, b]. This is the Uniform distribution. If you know a value lies between a and b but have no reason to prefer one sub-interval over another, the uniform distribution is the honest choice.
Examples: the position of a random point on a line segment, the phase angle of an unmodulated carrier wave, the round-off error when converting a real number to a fixed decimal representation.
The Normal (Gaussian) Distribution
No distribution in all of statistics is more important than the Normal distribution — the iconic bell curve. It appears everywhere: heights of people in a population, measurement errors in scientific instruments, returns on a financial asset over short periods, noise in electronic circuits, and countless other phenomena.
The Normal distribution is characterized by two parameters: the mean μ (which controls the center) and the standard deviation σ (which controls the spread). Its bell-shaped PDF is symmetric around μ and extends in both directions to ±∞.
For any normal distribution N(μ, σ²):
• About 68% of values fall within μ ± 1σ
• About 95% of values fall within μ ± 2σ
• About 99.7% of values fall within μ ± 3σ
This rule lets you quickly estimate probabilities without looking up a table.
Why does the normal distribution appear so ubiquitously? The answer is the Central Limit Theorem (the subject of the next lesson): the sum of many independent, identically distributed random variables converges to a normal distribution, regardless of the underlying distribution, provided the variance is finite. Any measured quantity that results from the sum of many small, independent influences tends to be normally distributed.
The Standard Normal and Z-Scores
The standard normal distribution Z ~ N(0, 1) has mean 0 and variance 1. It is the reference normal distribution from which all other normal distributions are derived. Any normal random variable X ~ N(μ, σ²) can be standardized to Z by the transformation:
Z-scores are invaluable for comparing observations from different normal distributions. A student who scores 75 on an exam with mean 65 and σ = 10 has Z = 1.0. A student who scores 120 on an IQ test with mean 100 and σ = 15 has Z = 1.33. The second student is further above average in their respective population, despite the lower absolute score on the exam scale.
The cumulative distribution function of Z, Φ(z) = P(Z ≤ z), gives the area under the standard normal PDF from −∞ to z. Tables of Φ(z) are available for any z, and modern software evaluates it directly. For any normal random variable, P(X ≤ x) = Φ((x − μ) / σ).
The Exponential Distribution
Where the Geometric distribution models the number of discrete trials until the first success, the Exponential distribution is its continuous analogue: it models the waiting time until the first event in a Poisson process. If events arrive at rate λ per unit time, the waiting time until the first event is Exponential(λ).
Applications abound: the lifetime of an electronic component before failure, the time between consecutive customer arrivals at a service desk, the gap between earthquakes in a seismically active region, the time until a radioactive atom decays.
Like the Geometric distribution, the Exponential distribution is memoryless: P(X > s + t | X > s) = P(X > t). If a device has been running for s hours without failure, the distribution of remaining life is the same as it was when new. Among continuous distributions, the Exponential is the unique memoryless distribution — a mathematical statement that the component is not aging or fatiguing.
The Gamma Distribution
The Gamma distribution generalizes the Exponential the same way the Negative Binomial generalizes the Geometric: it models the waiting time until the k-th event in a Poisson process. It is parameterized by a shape parameter k > 0 and a rate parameter λ > 0.
The Gamma family is remarkably flexible. With integer k, it is a sum of k independent Exponential(λ) variables. The special case k = n/2 and λ = 1/2 yields the chi-squared distribution with n degrees of freedom, ubiquitous in hypothesis testing and confidence interval construction.
The Beta Distribution
The Beta distribution lives on the interval [0, 1] and is the natural distribution for modeling probabilities of probabilities — or more generally, any quantity that must lie between 0 and 1, such as proportions, success rates, and fractions.
It is the conjugate prior for the Binomial distribution in Bayesian statistics: if your prior belief about a success probability p is Beta(α, β), and you observe k successes in n trials, the posterior is Beta(α + k, β + n − k). This clean updating rule makes the Beta indispensable in Bayesian A/B testing.
The shape of the Beta distribution is highly flexible. Beta(1, 1) is the Uniform distribution. Beta(α, β) with α < 1 and β < 1 produces a U-shape. With α = β > 1 the distribution is symmetric and bell-shaped. With α > β the distribution skews left; with α < β it skews right. This flexibility makes Beta an excellent model for expert opinion, Bayesian priors, and fractions from empirical data.
The Log-Normal Distribution
If X is normally distributed, then Y = eX follows a log-normal distribution. Equivalently, Y is log-normal if log(Y) is normal. The log-normal distribution is right-skewed and lives on (0, ∞), making it appropriate for quantities that are always positive and arise from multiplicative processes.
Examples: stock prices (daily percentage returns are approximately normal, so prices multiply → log-normal), incomes, size of biological organisms, droplet sizes in aerosols, file sizes on a computer. The multiplicative equivalent of the Central Limit Theorem explains why log-normal distributions appear so often in nature and economics.
Choosing the Right Continuous Distribution
Each continuous distribution encodes a generative story. Identifying the right story for your data immediately points you to the right model.
Equal probability over a bounded interval → Uniform(a, b)
Sum of many small independent influences; bell-shaped → Normal(μ, σ²)
Waiting time for first event in Poisson process → Exponential(λ)
Waiting time for k-th event in Poisson process → Gamma(k, λ)
Modeling a probability or proportion in [0, 1] → Beta(α, β)
Positive, right-skewed, from multiplicative process → Log-normal(μ, σ²)
Practical diagnostics: plot a histogram of your data and check the shape. If it is symmetric and bell-shaped, try Normal. If it is right-skewed and bounded below by zero, try Exponential, Gamma, or Log-normal — take the logarithm and check if that looks normal. If the data lies in (0, 1), try Beta. If you have no prior information about an interval quantity, use Uniform as the maximally uninformative model.
- Uniform(a, b): constant density on [a, b]. Mean = (a + b)/2. The maximally uninformative continuous distribution over a known interval.
- Normal(μ, σ²): the bell curve. Symmetric, characterized by mean and variance. The 68-95-99.7 rule: 68%, 95%, 99.7% of mass within 1, 2, 3 standard deviations of the mean.
- Standard normal Z ~ N(0,1): any normal variable standardizes to Z via Z = (X − μ)/σ. The CDF Φ(z) gives tail probabilities.
- Exponential(λ): waiting time for first Poisson event. Mean = 1/λ. The unique continuous memoryless distribution. CDF = 1 − e−λx.
- Gamma(k, λ): waiting time for k-th Poisson event. Generalizes Exponential. Chi-squared is a special case. Highly flexible shape.
- Beta(α, β): models probabilities and proportions on [0,1]. Conjugate prior for Binomial. Uniform is Beta(1,1).
- Log-normal(μ, σ²): Y is log-normal iff log(Y) is normal. Right-skewed, positive, arises from multiplicative processes. Models stock prices, incomes, organism sizes.