Reading
Stories Mode

Common Discrete Distributions

~18 min read Lesson 2 of 4 in Module 4

A Taxonomy of Counting Experiments

Nature is full of counting experiments. How many customers arrive in an hour? How many trials until the first defect? How many earthquakes strike a region each year? Each of these situations has a natural mathematical model — a named discrete distribution whose properties are well understood and whose parameters have direct real-world meaning.

Learning these distributions is not memorization for its own sake. Each one encodes a story about how data is generated. When you recognize the story behind your data, you immediately know the right model to use, its expected value, its variance, and how to compute any probability you need.

In this lesson we study five fundamental discrete distributions: Bernoulli, Binomial, Poisson, Geometric, and Negative Binomial. Together they cover an enormous range of real-world counting problems.

The Bernoulli Distribution

The simplest possible random variable takes only two values: 1 (success) with probability p, and 0 (failure) with probability 1 − p. This is the Bernoulli distribution, named after Swiss mathematician Jacob Bernoulli.

Every binary outcome in statistics — a coin flip, a yes/no survey response, whether a manufactured part passes inspection, whether a patient recovers — is a Bernoulli trial. The parameter p is the success probability, and it completely determines the distribution.

Bernoulli PMF
P(X = k) = p^k (1-p)^{1-k}, \quad k \in \{0, 1\}
P(X = 1) = p and P(X = 0) = 1 − p. The single parameter p ∈ (0, 1) is the probability of success.
Mean
E[X] = p
The expected value equals the success probability
Variance
\text{Var}(X) = p(1-p)
Variance is maximized when p = 0.5

The variance p(1 − p) reaches its maximum of 0.25 at p = 0.5, and approaches zero as p approaches 0 or 1. This makes intuitive sense: a very biased coin (nearly always heads) is more predictable, hence lower variance, than a fair coin.

The Binomial Distribution

Now suppose you repeat a Bernoulli trial n times, independently, with the same success probability p on each trial. The Binomial distribution models the total number of successes across all n trials.

The binomial distribution is arguably the most important discrete distribution in all of statistics. It underlies clinical trials (how many patients respond to treatment?), quality control (how many items are defective in a batch?), election polling (how many voters favor a candidate?), and A/B testing (how many users click the new button?).

Binomial PMF
P(X = k) = \binom{n}{k} p^k (1-p)^{n-k}
The binomial coefficient C(n,k) counts the number of ways to choose which k of the n trials are successes.
Binomial Story

Setup: n independent trials, each with success probability p.

Random variable X: total number of successes, X ∈ {0, 1, 2, …, n}.

Key assumption: trials are independent and all have the same p.

Mean
E[X] = np
Average number of successes in n trials
Variance
\text{Var}(X) = np(1-p)
Spread of the success count around the mean

The mean np follows immediately from linearity of expectation: X is the sum of n independent Bernoulli(p) variables, each with mean p, so E[X] = np. Similarly Var(X) = n × p(1 − p) by the independence of the trials.

Example: suppose 30% of customers who receive a promotional email make a purchase. If we send 200 emails, the expected number of purchases is E[X] = 200 × 0.3 = 60, with standard deviation σ = √(200 × 0.3 × 0.7) ≈ 6.5.

The Poisson Distribution

Some random variables count events that occur in a fixed period of time or space — phone calls arriving at a call center per hour, accidents on a highway per month, typos per page of manuscript. When the events are rare relative to the opportunities for them to occur, and when they occur independently, the Poisson distribution is the natural model.

The Poisson distribution has a single parameter λ (lambda), the rate or average number of events in the interval. Remarkably, both the mean and variance equal λ.

Poisson PMF
P(X = k) = \frac{e^{-\lambda}\,\lambda^k}{k!}
λ > 0 is the average rate of events. The support is k = 0, 1, 2, … (no upper limit).
Mean
E[X] = \lambda
Mean equals the rate parameter λ
Variance
\text{Var}(X) = \lambda
Variance also equals λ — a unique property

The Poisson distribution arises as a limit of the Binomial when n is very large and p is very small, with np = λ held fixed. This is the Poisson approximation to the binomial: if n ≥ 100 and p ≤ 0.01, the Poisson(λ = np) is an excellent approximation.

Poisson Conditions

Events occur independently of one another.

The rate λ is constant over the observation period.

Two events cannot occur at the exact same instant (no simultaneous events).

The Geometric Distribution

Instead of counting successes in a fixed number of trials, suppose you ask: how many trials until the first success? This waiting-time question leads to the Geometric distribution.

Imagine repeatedly rolling a die until you get a six. How many rolls will it take? Each roll independently has probability p = 1/6 of success. The Geometric distribution models the number of trials needed, including the successful one.

Geometric PMF
P(X = k) = (1-p)^{k-1}\,p
k = 1, 2, 3, … is the trial number on which the first success occurs. The first k − 1 trials are failures, the k-th is a success.
Mean
E[X] = \dfrac{1}{p}
Average trials until first success
Variance
\text{Var}(X) = \dfrac{1-p}{p^2}
Spread of the waiting time

The Geometric distribution has a remarkable property called memorylessness: given that you have already failed k times, the distribution of additional trials needed is the same as if you were starting fresh. In other words, the past failures give you no information about when the next success will come. Among discrete distributions, the Geometric is the unique memoryless distribution.

For p = 1/6 (rolling a six), E[X] = 6 — on average you need six rolls to get a six. For p = 0.01 (a 1-in-100 rare event), E[X] = 100 — you need about 100 attempts on average.

The Negative Binomial Distribution

The Geometric asks: how many trials until the first success? The Negative Binomial distribution generalizes this to: how many trials until the r-th success? It is the natural model when you need to wait for multiple successes.

Real examples: how many manufacturing attempts before r acceptable units are produced? How many patients must be recruited before r respond to a treatment? How many emails must be sent before r conversions occur?

Negative Binomial PMF
P(X = k) = \binom{k-1}{r-1} p^r (1-p)^{k-r}
k ≥ r is the trial on which the r-th success occurs. C(k−1, r−1) counts arrangements of the first r−1 successes among the first k−1 trials.
Mean
E[X] = \dfrac{r}{p}
Average trials to achieve r successes
Variance
\text{Var}(X) = \dfrac{r(1-p)}{p^2}
Spread of the waiting time for r successes

The Negative Binomial with r = 1 reduces to the Geometric, as expected. It can also be viewed as a sum of r independent Geometric(p) random variables: if you split the waiting time for the r-th success into r successive waits for each individual success, each waiting time is Geometric(p). Linearity then gives the mean r/p and the variance r(1 − p)/p² immediately.

Choosing the Right Distribution

Each of the five distributions we have studied corresponds to a specific generative story. Identifying which story matches your data is the key skill in applied statistics.

Distribution Decision Guide

Single binary trial (success/failure, one shot) → Bernoulli(p)

Fixed number of independent trials, count successes → Binomial(n, p)

Count of rare, independent events in a fixed interval → Poisson(λ)

How many trials until the first success? → Geometric(p)

How many trials until the r-th success? → Negative Binomial(r, p)

A common diagnostic check: if the mean and variance of your data are roughly equal, Poisson is a strong candidate. If the variance is much larger than the mean (overdispersion), the Negative Binomial is often a better fit. If the variance is much smaller than the mean (underdispersion), consider checking whether the independence assumption holds.

Also watch for parameter estimation: for the Binomial, estimate p = (observed successes) / n. For the Poisson, estimate λ = sample mean. For the Geometric, estimate p = 1 / sample mean. These are the maximum likelihood estimates, obtained by matching the theoretical mean to the observed mean.

Key Takeaways
  • Bernoulli(p): single binary trial. Mean = p, Var = p(1 − p). The atom from which all other discrete distributions are built.
  • Binomial(n, p): count of successes in n independent Bernoulli trials. Mean = np, Var = np(1 − p).
  • Poisson(λ): count of rare, independent events in a fixed interval. Mean = Var = λ. Limit of Binomial as n → ∞, p → 0, np = λ.
  • Geometric(p): trials until first success. Mean = 1/p, Var = (1 − p)/p². The unique memoryless discrete distribution.
  • Negative Binomial(r, p): trials until r-th success. Mean = r/p, Var = r(1 − p)/p². Sum of r independent Geometric(p) variables.
  • Recognizing the story behind your data immediately determines the correct model, its parameters, and its key quantities.
  • If mean ≈ variance ⇒ consider Poisson. If variance ≫ mean ⇒ consider Negative Binomial. If discrete + fixed n ⇒ consider Binomial.
Previous Random Variables Module Overview Next Common Continuous Distributions