Reading
Stories Mode

Common Continuous Distributions

~20 min read Lesson 3 of 4 in Module 4

From Counting to Measuring

In the previous lesson we studied distributions for counting things — successes in trials, arrivals in an interval, trials until a success. Now we shift to continuous random variables, which model quantities we measure rather than count: height, time, temperature, price, distance.

For a continuous random variable X, we no longer assign probability to individual values — the probability that X equals exactly 3.14159 is zero. Instead, probability is assigned to intervals via the probability density function (PDF). The PDF f(x) satisfies f(x) ≥ 0 and ∫ f(x) dx = 1, and P(a ≤ X ≤ b) = ∫ab f(x) dx.

In this lesson we survey seven fundamental continuous distributions: Uniform, Normal (Gaussian), Exponential, Gamma, Beta, and Log-normal. Together they underpin an enormous range of engineering, scientific, and data analysis problems.

The Uniform Distribution

The simplest continuous distribution assigns equal probability density to every point in an interval [a, b]. This is the Uniform distribution. If you know a value lies between a and b but have no reason to prefer one sub-interval over another, the uniform distribution is the honest choice.

Examples: the position of a random point on a line segment, the phase angle of an unmodulated carrier wave, the round-off error when converting a real number to a fixed decimal representation.

Uniform PDF
f(x) = \dfrac{1}{b-a}, \quad a \le x \le b
f(x) is constant over [a, b] and zero outside. The constant value 1/(b − a) ensures the total area equals 1.
Mean
E[X] = \dfrac{a+b}{2}
The midpoint of the interval
Variance
\text{Var}(X) = \dfrac{(b-a)^2}{12}
Spread increases with the interval width

The Normal (Gaussian) Distribution

No distribution in all of statistics is more important than the Normal distribution — the iconic bell curve. It appears everywhere: heights of people in a population, measurement errors in scientific instruments, returns on a financial asset over short periods, noise in electronic circuits, and countless other phenomena.

The Normal distribution is characterized by two parameters: the mean μ (which controls the center) and the standard deviation σ (which controls the spread). Its bell-shaped PDF is symmetric around μ and extends in both directions to ±∞.

Normal PDF
f(x) = \dfrac{1}{\sigma\sqrt{2\pi}}\,e^{-\frac{(x-\mu)^2}{2\sigma^2}}
μ is the mean (location) and σ > 0 is the standard deviation (scale). The factor 1/(σ√2π) normalizes the area to 1.
The 68-95-99.7 Rule

For any normal distribution N(μ, σ²):

• About 68% of values fall within μ ± 1σ

• About 95% of values fall within μ ± 2σ

• About 99.7% of values fall within μ ± 3σ

This rule lets you quickly estimate probabilities without looking up a table.

Why does the normal distribution appear so ubiquitously? The answer is the Central Limit Theorem (the subject of the next lesson): the sum of many independent, identically distributed random variables converges to a normal distribution, regardless of the underlying distribution, provided the variance is finite. Any measured quantity that results from the sum of many small, independent influences tends to be normally distributed.

The Standard Normal and Z-Scores

The standard normal distribution Z ~ N(0, 1) has mean 0 and variance 1. It is the reference normal distribution from which all other normal distributions are derived. Any normal random variable X ~ N(μ, σ²) can be standardized to Z by the transformation:

Z-Score
Z = \dfrac{X - \mu}{\sigma}
Z measures how many standard deviations X is from its mean. Positive Z means above average; negative Z means below average.

Z-scores are invaluable for comparing observations from different normal distributions. A student who scores 75 on an exam with mean 65 and σ = 10 has Z = 1.0. A student who scores 120 on an IQ test with mean 100 and σ = 15 has Z = 1.33. The second student is further above average in their respective population, despite the lower absolute score on the exam scale.

The cumulative distribution function of Z, Φ(z) = P(Z ≤ z), gives the area under the standard normal PDF from −∞ to z. Tables of Φ(z) are available for any z, and modern software evaluates it directly. For any normal random variable, P(X ≤ x) = Φ((x − μ) / σ).

The Exponential Distribution

Where the Geometric distribution models the number of discrete trials until the first success, the Exponential distribution is its continuous analogue: it models the waiting time until the first event in a Poisson process. If events arrive at rate λ per unit time, the waiting time until the first event is Exponential(λ).

Applications abound: the lifetime of an electronic component before failure, the time between consecutive customer arrivals at a service desk, the gap between earthquakes in a seismically active region, the time until a radioactive atom decays.

Exponential PDF
f(x) = \lambda\,e^{-\lambda x}, \quad x \ge 0
λ > 0 is the rate parameter. The PDF starts at λ when x = 0 and decays exponentially. The CDF is F(x) = 1 − e−λx.
Mean
E[X] = \dfrac{1}{\lambda}
Average waiting time is 1/λ
Variance
\text{Var}(X) = \dfrac{1}{\lambda^2}
Standard deviation also equals 1/λ

Like the Geometric distribution, the Exponential distribution is memoryless: P(X > s + t | X > s) = P(X > t). If a device has been running for s hours without failure, the distribution of remaining life is the same as it was when new. Among continuous distributions, the Exponential is the unique memoryless distribution — a mathematical statement that the component is not aging or fatiguing.

The Gamma Distribution

The Gamma distribution generalizes the Exponential the same way the Negative Binomial generalizes the Geometric: it models the waiting time until the k-th event in a Poisson process. It is parameterized by a shape parameter k > 0 and a rate parameter λ > 0.

Gamma PDF
f(x) = \dfrac{\lambda^k x^{k-1} e^{-\lambda x}}{\Gamma(k)}, \quad x \ge 0
Γ(k) is the gamma function, generalizing the factorial: Γ(k) = (k − 1)! for positive integers k. When k = 1, the Gamma reduces to the Exponential.
Mean
E[X] = \dfrac{k}{\lambda}
Average time for k events to occur
Variance
\text{Var}(X) = \dfrac{k}{\lambda^2}
Spread of the waiting time for k events

The Gamma family is remarkably flexible. With integer k, it is a sum of k independent Exponential(λ) variables. The special case k = n/2 and λ = 1/2 yields the chi-squared distribution with n degrees of freedom, ubiquitous in hypothesis testing and confidence interval construction.

The Beta Distribution

The Beta distribution lives on the interval [0, 1] and is the natural distribution for modeling probabilities of probabilities — or more generally, any quantity that must lie between 0 and 1, such as proportions, success rates, and fractions.

It is the conjugate prior for the Binomial distribution in Bayesian statistics: if your prior belief about a success probability p is Beta(α, β), and you observe k successes in n trials, the posterior is Beta(α + k, β + n − k). This clean updating rule makes the Beta indispensable in Bayesian A/B testing.

Beta PDF
f(x) = \dfrac{x^{\alpha-1}(1-x)^{\beta-1}}{B(\alpha,\beta)}, \quad 0 \le x \le 1
α > 0 and β > 0 are shape parameters. B(α, β) = Γ(α)Γ(β)/Γ(α+β) is the beta function, ensuring normalization.

The shape of the Beta distribution is highly flexible. Beta(1, 1) is the Uniform distribution. Beta(α, β) with α < 1 and β < 1 produces a U-shape. With α = β > 1 the distribution is symmetric and bell-shaped. With α > β the distribution skews left; with α < β it skews right. This flexibility makes Beta an excellent model for expert opinion, Bayesian priors, and fractions from empirical data.

The Log-Normal Distribution

If X is normally distributed, then Y = eX follows a log-normal distribution. Equivalently, Y is log-normal if log(Y) is normal. The log-normal distribution is right-skewed and lives on (0, ∞), making it appropriate for quantities that are always positive and arise from multiplicative processes.

Examples: stock prices (daily percentage returns are approximately normal, so prices multiply → log-normal), incomes, size of biological organisms, droplet sizes in aerosols, file sizes on a computer. The multiplicative equivalent of the Central Limit Theorem explains why log-normal distributions appear so often in nature and economics.

Mean
E[Y] = e^{\mu + \sigma^2/2}
μ and σ are the mean and std of log(Y)
Variance
\text{Var}(Y) = (e^{\sigma^2}-1)\,e^{2\mu+\sigma^2}
Variance grows quickly with σ

Choosing the Right Continuous Distribution

Each continuous distribution encodes a generative story. Identifying the right story for your data immediately points you to the right model.

Distribution Decision Guide

Equal probability over a bounded interval → Uniform(a, b)

Sum of many small independent influences; bell-shaped → Normal(μ, σ²)

Waiting time for first event in Poisson process → Exponential(λ)

Waiting time for k-th event in Poisson process → Gamma(k, λ)

Modeling a probability or proportion in [0, 1] → Beta(α, β)

Positive, right-skewed, from multiplicative process → Log-normal(μ, σ²)

Practical diagnostics: plot a histogram of your data and check the shape. If it is symmetric and bell-shaped, try Normal. If it is right-skewed and bounded below by zero, try Exponential, Gamma, or Log-normal — take the logarithm and check if that looks normal. If the data lies in (0, 1), try Beta. If you have no prior information about an interval quantity, use Uniform as the maximally uninformative model.

Key Takeaways
  • Uniform(a, b): constant density on [a, b]. Mean = (a + b)/2. The maximally uninformative continuous distribution over a known interval.
  • Normal(μ, σ²): the bell curve. Symmetric, characterized by mean and variance. The 68-95-99.7 rule: 68%, 95%, 99.7% of mass within 1, 2, 3 standard deviations of the mean.
  • Standard normal Z ~ N(0,1): any normal variable standardizes to Z via Z = (X − μ)/σ. The CDF Φ(z) gives tail probabilities.
  • Exponential(λ): waiting time for first Poisson event. Mean = 1/λ. The unique continuous memoryless distribution. CDF = 1 − e−λx.
  • Gamma(k, λ): waiting time for k-th Poisson event. Generalizes Exponential. Chi-squared is a special case. Highly flexible shape.
  • Beta(α, β): models probabilities and proportions on [0,1]. Conjugate prior for Binomial. Uniform is Beta(1,1).
  • Log-normal(μ, σ²): Y is log-normal iff log(Y) is normal. Right-skewed, positive, arises from multiplicative processes. Models stock prices, incomes, organism sizes.
Previous Common Discrete Distributions Module Overview Next The Central Limit Theorem