Chance is Everywhere
When you flip a coin, roll a die, check the weather forecast, or receive a medical test result, you are dealing with uncertainty. Probability is the mathematical language we use to reason about that uncertainty rigorously — to assign numbers to how likely events are, and to update those numbers as new information arrives.
The study of probability is the foundation on which all of statistics rests. Before we can estimate parameters, test hypotheses, or build models, we need a consistent set of rules for talking about randomness. Those rules are what this lesson is about.
Statistics without probability is just description. Probability lets us move from “here is what happened” to “here is what we can infer about the process that generated the data” — and that is where the real power lies.
Probability is the calculus of uncertainty.Experiments, Outcomes, and Sample Spaces
Every probabilistic setting starts with a random experiment: any process whose outcome cannot be predicted with certainty in advance. Flipping a fair coin, rolling a six-sided die, measuring tomorrow’s temperature, and randomly selecting a person from a population are all examples.
The complete list of all possible outcomes of a random experiment is called the sample space, usually denoted Ω (omega). For a coin flip, Ω = {H, T}. For a single die roll, Ω = {1, 2, 3, 4, 5, 6}. For tossing two coins, Ω = {HH, HT, TH, TT}.
An event is any subset of the sample space — any collection of outcomes we care about. “Rolling an even number” corresponds to the event A = {2, 4, 6}. “Getting at least one head in two coin flips” corresponds to B = {HH, HT, TH}.
Random experiment — a process with an uncertain outcome
Sample space Ω — the set of all possible outcomes
Event — a subset of the sample space
Elementary event — a single outcome (a subset of size 1)
Event Operations
Events are sets, so all the operations of set theory apply to them. Let A and B be two events in the same sample space Ω.
The union A ∪ B contains all outcomes in A, in B, or in both. It corresponds to the event “A or B (or both) occur.” For rolling a die with A = {2, 4, 6} (even) and B = {1, 2, 3} (less than 4), A ∪ B = {1, 2, 3, 4, 6}.
The intersection A ∩ B contains outcomes in both A and B simultaneously. It corresponds to “A and B both occur.” Using the same example, A ∩ B = {2}.
The complement Ac (also written  or A′) contains all outcomes in Ω that are not in A. It corresponds to “A does not occur.” For A = {2, 4, 6}, Ac = {1, 3, 5}.
Two events are mutually exclusive (or disjoint) if A ∩ B = ∅ — they cannot both occur on the same trial.
Three Interpretations of Probability
Mathematicians agree on the axioms of probability (more on those shortly), but the philosophical meaning of a probability statement has been debated for centuries. There are three major schools of thought.
The classical interpretation applies when all outcomes are equally likely. Probability is the ratio of favorable outcomes to total outcomes. Rolling a fair die, the probability of getting a 3 is 1/6 because exactly one of the six equally likely faces shows a 3.
The frequentist interpretation defines probability as the long-run relative frequency. If you flip a fair coin one million times, the fraction of heads converges to 0.5. Probability is not a property of a single trial but of an indefinitely repeated process.
The Bayesian interpretation treats probability as a degree of belief — a measure of confidence that an event will occur, given everything we currently know. Under this view, probability can be assigned to one-time events and updated as evidence accumulates. The probability that the defendant is guilty, that a new drug works, or that it will rain tomorrow can all be given Bayesian probabilities.
Classical: ratio of favorable to total equally-likely outcomes
Frequentist: long-run relative frequency over many repetitions
Bayesian: subjective degree of belief, updated by evidence
Kolmogorov’s Axioms
Regardless of which interpretation you prefer, all three agree on the mathematical rules. In 1933, Andrei Kolmogorov placed probability theory on a rigorous axiomatic foundation. His three axioms are simple but sufficient to derive every result in probability theory.
Axiom 1 (Non-negativity): For any event A, P(A) ≥ 0.
Axiom 2 (Normalization): P(Ω) = 1. Something must happen.
Axiom 3 (Additivity): For mutually exclusive events A and B, P(A ∪ B) = P(A) + P(B).
Everything else in probability follows from these three axioms. Probabilities are always between 0 and 1. The probability of the impossible event is 0. Complementary events sum to 1. All of these can be proved from the axioms — they are not additional assumptions.
Core Probability Rules
From Kolmogorov’s axioms, two essential rules follow immediately.
The complement rule states that the probability of an event not happening equals one minus the probability that it does happen. This is often the easiest way to compute “at least one” probabilities: find the probability of the complement (none) and subtract from 1.
The addition rule gives the probability of the union of two events. If A and B can overlap, we must subtract the double-counted intersection.
Counting: Permutations and Combinations
For classical probability, computing P(A) = |A| / |Ω| requires knowing how many outcomes are in each set. This is where counting methods — combinatorics — become essential.
A permutation is an ordered arrangement of items. If you have n distinct objects and choose k of them in order, the number of permutations is:
A combination is an unordered selection — a subset. If order does not matter (we only care which items are chosen, not their sequence), the count is smaller:
Combinations appear constantly in probability. The probability of being dealt a specific poker hand, of choosing a committee, of any scenario where order is irrelevant — all reduce to computing binomial coefficients C(n, k).
Permutations: order matters. The PIN 1-2-3-4 is different from 4-3-2-1.
Combinations: order does not matter. The committee {Alice, Bob, Carol} is the same regardless of the order they were chosen.
C(n, k) = P(n, k) / k! — divide out the k! redundant orderingsPutting It All Together
Probability gives us a language: sample spaces define all possible outcomes, events are the things we care about, and probability measures assign numbers between 0 and 1 consistently. The axioms guarantee that our assignments do not contradict each other. The complement and addition rules let us compute probabilities for complex events from simpler ones. Counting methods let us enumerate possibilities when the classical setting applies.
In the next lesson, we will extend this framework to conditional probability — what happens to the probability of A when we learn that B has occurred — and to one of the most powerful tools in all of statistics: Bayes’ theorem.
- A random experiment has a sample space Ω of all possible outcomes; an event is any subset of Ω.
- Events combine via union (∪), intersection (∩), and complement (·c) — exactly like sets.
- Three interpretations of probability — classical, frequentist, Bayesian — agree on the axioms but differ on meaning.
- Kolmogorov’s three axioms (non-negativity, normalization, additivity) are the foundation of all probability theory.
- The complement rule and the general addition rule are the two most-used immediate consequences of the axioms.
- Permutations count ordered arrangements; combinations count unordered subsets — both are essential for classical probability calculations.