A Taxonomy of Counting Experiments
Nature is full of counting experiments. How many customers arrive in an hour? How many trials until the first defect? How many earthquakes strike a region each year? Each of these situations has a natural mathematical model — a named discrete distribution whose properties are well understood and whose parameters have direct real-world meaning.
Learning these distributions is not memorization for its own sake. Each one encodes a story about how data is generated. When you recognize the story behind your data, you immediately know the right model to use, its expected value, its variance, and how to compute any probability you need.
In this lesson we study five fundamental discrete distributions: Bernoulli, Binomial, Poisson, Geometric, and Negative Binomial. Together they cover an enormous range of real-world counting problems.
The Bernoulli Distribution
The simplest possible random variable takes only two values: 1 (success) with probability p, and 0 (failure) with probability 1 − p. This is the Bernoulli distribution, named after Swiss mathematician Jacob Bernoulli.
Every binary outcome in statistics — a coin flip, a yes/no survey response, whether a manufactured part passes inspection, whether a patient recovers — is a Bernoulli trial. The parameter p is the success probability, and it completely determines the distribution.
The variance p(1 − p) reaches its maximum of 0.25 at p = 0.5, and approaches zero as p approaches 0 or 1. This makes intuitive sense: a very biased coin (nearly always heads) is more predictable, hence lower variance, than a fair coin.
The Binomial Distribution
Now suppose you repeat a Bernoulli trial n times, independently, with the same success probability p on each trial. The Binomial distribution models the total number of successes across all n trials.
The binomial distribution is arguably the most important discrete distribution in all of statistics. It underlies clinical trials (how many patients respond to treatment?), quality control (how many items are defective in a batch?), election polling (how many voters favor a candidate?), and A/B testing (how many users click the new button?).
Setup: n independent trials, each with success probability p.
Random variable X: total number of successes, X ∈ {0, 1, 2, …, n}.
Key assumption: trials are independent and all have the same p.
The mean np follows immediately from linearity of expectation: X is the sum of n independent Bernoulli(p) variables, each with mean p, so E[X] = np. Similarly Var(X) = n × p(1 − p) by the independence of the trials.
Example: suppose 30% of customers who receive a promotional email make a purchase. If we send 200 emails, the expected number of purchases is E[X] = 200 × 0.3 = 60, with standard deviation σ = √(200 × 0.3 × 0.7) ≈ 6.5.
The Poisson Distribution
Some random variables count events that occur in a fixed period of time or space — phone calls arriving at a call center per hour, accidents on a highway per month, typos per page of manuscript. When the events are rare relative to the opportunities for them to occur, and when they occur independently, the Poisson distribution is the natural model.
The Poisson distribution has a single parameter λ (lambda), the rate or average number of events in the interval. Remarkably, both the mean and variance equal λ.
The Poisson distribution arises as a limit of the Binomial when n is very large and p is very small, with np = λ held fixed. This is the Poisson approximation to the binomial: if n ≥ 100 and p ≤ 0.01, the Poisson(λ = np) is an excellent approximation.
Events occur independently of one another.
The rate λ is constant over the observation period.
Two events cannot occur at the exact same instant (no simultaneous events).
The Geometric Distribution
Instead of counting successes in a fixed number of trials, suppose you ask: how many trials until the first success? This waiting-time question leads to the Geometric distribution.
Imagine repeatedly rolling a die until you get a six. How many rolls will it take? Each roll independently has probability p = 1/6 of success. The Geometric distribution models the number of trials needed, including the successful one.
The Geometric distribution has a remarkable property called memorylessness: given that you have already failed k times, the distribution of additional trials needed is the same as if you were starting fresh. In other words, the past failures give you no information about when the next success will come. Among discrete distributions, the Geometric is the unique memoryless distribution.
For p = 1/6 (rolling a six), E[X] = 6 — on average you need six rolls to get a six. For p = 0.01 (a 1-in-100 rare event), E[X] = 100 — you need about 100 attempts on average.
The Negative Binomial Distribution
The Geometric asks: how many trials until the first success? The Negative Binomial distribution generalizes this to: how many trials until the r-th success? It is the natural model when you need to wait for multiple successes.
Real examples: how many manufacturing attempts before r acceptable units are produced? How many patients must be recruited before r respond to a treatment? How many emails must be sent before r conversions occur?
The Negative Binomial with r = 1 reduces to the Geometric, as expected. It can also be viewed as a sum of r independent Geometric(p) random variables: if you split the waiting time for the r-th success into r successive waits for each individual success, each waiting time is Geometric(p). Linearity then gives the mean r/p and the variance r(1 − p)/p² immediately.
Choosing the Right Distribution
Each of the five distributions we have studied corresponds to a specific generative story. Identifying which story matches your data is the key skill in applied statistics.
Single binary trial (success/failure, one shot) → Bernoulli(p)
Fixed number of independent trials, count successes → Binomial(n, p)
Count of rare, independent events in a fixed interval → Poisson(λ)
How many trials until the first success? → Geometric(p)
How many trials until the r-th success? → Negative Binomial(r, p)
A common diagnostic check: if the mean and variance of your data are roughly equal, Poisson is a strong candidate. If the variance is much larger than the mean (overdispersion), the Negative Binomial is often a better fit. If the variance is much smaller than the mean (underdispersion), consider checking whether the independence assumption holds.
Also watch for parameter estimation: for the Binomial, estimate p = (observed successes) / n. For the Poisson, estimate λ = sample mean. For the Geometric, estimate p = 1 / sample mean. These are the maximum likelihood estimates, obtained by matching the theoretical mean to the observed mean.
- Bernoulli(p): single binary trial. Mean = p, Var = p(1 − p). The atom from which all other discrete distributions are built.
- Binomial(n, p): count of successes in n independent Bernoulli trials. Mean = np, Var = np(1 − p).
- Poisson(λ): count of rare, independent events in a fixed interval. Mean = Var = λ. Limit of Binomial as n → ∞, p → 0, np = λ.
- Geometric(p): trials until first success. Mean = 1/p, Var = (1 − p)/p². The unique memoryless discrete distribution.
- Negative Binomial(r, p): trials until r-th success. Mean = r/p, Var = r(1 − p)/p². Sum of r independent Geometric(p) variables.
- Recognizing the story behind your data immediately determines the correct model, its parameters, and its key quantities.
- If mean ≈ variance ⇒ consider Poisson. If variance ≫ mean ⇒ consider Negative Binomial. If discrete + fixed n ⇒ consider Binomial.