Reading
Stories Mode

Conditional Probability

~14 min read Lesson 2 of 3 in Module 3

When Context Changes Everything

Imagine you are told that a randomly chosen person tests positive for a rare disease. What is the probability that they actually have it? Now add one piece of context: the test has a 5% false-positive rate and the disease affects only 1 in 10,000 people. Suddenly the answer changes dramatically — perhaps the person is far more likely to be a false positive than a true case.

This is the essential insight of conditional probability: knowing something about the world changes the probability of everything else. It is the mathematical formalization of the phrase “given that…” and it underlies Bayes’ theorem, the law of total probability, and a vast swath of statistical reasoning.

The Core Idea

Unconditional probability asks: “How likely is A?” Conditional probability asks: “How likely is A, now that we know B has occurred?” Learning B restricts the sample space, and the conditional probability reflects that restriction.

P(A | B) — the probability of A given B

Defining P(A | B)

The conditional probability of event A given event B is defined as the probability that both A and B occur, divided by the probability that B occurs. Intuitively, knowing that B happened restricts our attention to outcomes inside B; within that restricted world, we ask what fraction also falls in A.

Conditional Probability
P(A \mid B) = \dfrac{P(A \cap B)}{P(B)}
Defined whenever P(B) > 0. If B is impossible, conditioning on B is undefined.

To see why this formula makes sense, consider a fair die. The sample space is {1, 2, 3, 4, 5, 6}. Let A = {6} (rolling a 6) and B = {4, 5, 6} (rolling more than 3). Unconditionally, P(A) = 1/6. But given that B occurred, we know the outcome is in {4, 5, 6}, and only one of those three is a 6. So P(A|B) = 1/3, which matches the formula: P(A ∩ B) = P({6}) = 1/6, P(B) = 1/2, and (1/6)/(1/2) = 1/3.

The formula also clarifies something important: conditioning on B is equivalent to renormalizing probabilities so that B becomes the new “certain event.” Every probability we knew is scaled by 1/P(B) — keeping relative proportions within B intact.

The Multiplication Rule

Rearranging the definition of conditional probability gives the multiplication rule, one of the most-used identities in probability. It allows us to compute the probability of two events both occurring by decomposing it into a chain of simpler probabilities.

Multiplication Rule
P(A \cap B) = P(A \mid B) \cdot P(B)
Equivalently P(A) · P(B|A) by symmetry — which factorization you use depends on which conditional probability is easier to evaluate

The rule extends naturally to chains of three or more events. The probability that A, B, and C all occur equals P(A) · P(B|A) · P(C|A ∩ B). This is the chain rule of probability, and it is the basis for probabilistic graphical models and Bayesian networks.

Worked Example

An urn contains 3 red and 5 blue balls. Two balls are drawn without replacement. What is the probability that both are red?

P(1st red) = 3/8. Given the first was red, 2 reds remain among 7: P(2nd red | 1st red) = 2/7.

P(both red) = (3/8) · (2/7) = 6/56 = 3/28 ≈ 0.107

Statistical Independence

Two events A and B are statistically independent if knowing that B occurred does not change the probability of A. Formally, A and B are independent if and only if P(A|B) = P(A). Substituting this into the multiplication rule gives an equivalent condition that is often more convenient to check.

Independence Condition
P(A \cap B) = P(A) \cdot P(B)
A and B are independent if and only if their joint probability factors into the product of their marginal probabilities

Independence is a strong assumption. In the urn example above, the two draws were dependent because removing one ball changed the composition for the second draw. If we draw with replacement, they become independent. In real data, independence is rarely guaranteed and must be verified or assumed carefully.

Note the difference between mutually exclusive and independent. Mutually exclusive events (A ∩ B = ∅) cannot both occur, so knowing A happened means B definitely did not — they are as dependent as possible. Independent events can both occur; it is just that one carries no information about the other.

Common Confusion

Mutually exclusive ≠ independent. If A and B are mutually exclusive and both have positive probability, they are maximally dependent: P(A|B) = 0 ≠ P(A).

Independent ≠ uncorrelated in general, though for normal distributions uncorrelated does imply independence.

The Law of Total Probability

Suppose the sample space Ω can be partitioned into mutually exclusive, exhaustive events B1, B2, …, Bn (they do not overlap and together cover everything). Then the probability of any event A can be computed by averaging over all the partitions, weighting each by how likely that partition is.

Law of Total Probability
P(A) = \sum_{i=1}^{n} P(A \mid B_i)\, P(B_i)
B1, …, Bn must be a partition of Ω: pairwise disjoint and collectively exhaustive

This law is indispensable for computing P(A) when A’s probability is easy to compute in each subgroup but not overall. It is also the denominator in Bayes’ theorem.

Example: Factory Quality

Factory A produces 60% of widgets (2% defective). Factory B produces 40% (5% defective). What fraction of all widgets is defective?

P(defect) = 0.6 · 0.02 + 0.4 · 0.05 = 0.012 + 0.020 = 0.032

Bayes’ Theorem

Bayes’ theorem answers a fundamental question: given that we have observed evidence B, how should we update our belief about hypothesis A? It inverts the direction of conditioning — instead of P(B|A) (likelihood of the evidence given the hypothesis), it gives us P(A|B) (probability of the hypothesis given the evidence).

Bayes’ Theorem
P(A \mid B) = \dfrac{P(B \mid A)\, P(A)}{P(B)}
P(A) is the prior, P(B|A) is the likelihood, P(B) is the marginal likelihood, and P(A|B) is the posterior

Combining Bayes’ theorem with the law of total probability gives the most general form. If {A, Ac} partitions Ω, then P(B) = P(B|A) · P(A) + P(B|Ac) · P(Ac), and substituting this into Bayes’ formula makes it fully computable from quantities we typically know.

The language of Bayesian reasoning — prior, likelihood, posterior — is central to a whole school of statistical inference. For now, understand Bayes’ theorem as a precise formula for updating probabilities in light of new evidence.

Tree Diagrams

A tree diagram is a visual tool for organizing conditional probability calculations. Each branch represents an event, labeled with its probability; the probability of any path through the tree is the product of probabilities along that path (the multiplication rule in action).

To use a tree diagram: (1) identify the first-stage events and place them on the first level of branches; (2) for each first-stage outcome, identify the second-stage conditional events and draw their branches; (3) multiply along each path to get joint probabilities; (4) add paths that lead to the same final outcome to compute marginal probabilities.

Reading a Tree

Along a path: multiply. The probability of reaching a leaf is the product of all branch probabilities on the path from root to leaf.

Across paths: add. The probability of a final outcome is the sum over all paths that lead to it.

P(leaf) = P(branch 1) · P(branch 2 | branch 1) · …

P(A|B) ≠ P(B|A): A Costly Mistake

One of the most common and consequential errors in probabilistic reasoning is confusing P(A|B) with P(B|A). They can be vastly different, and mixing them up leads to faulty conclusions in medicine, law, and science.

In medicine: P(positive test | disease) is the test’s sensitivity — typically high for good tests. But P(disease | positive test) is the positive predictive value, which depends critically on disease prevalence. For a rare disease, even a sensitive test will produce mostly false positives.

In law: the prosecutor’s fallacy conflates P(evidence | innocent) with P(innocent | evidence). A DNA match occurring in 1 in a million innocent people is not the same as a 1-in-a-million probability of innocence — the prior probability of guilt, determined by all other evidence, matters enormously.

The Transpose Error

P(A|B) = P(B|A) only if P(A) = P(B). In general they are related by Bayes’ theorem: P(A|B) = P(B|A) · P(A) / P(B). Whenever P(A) ≠ P(B), the two conditional probabilities will differ — sometimes by orders of magnitude.

Putting It All Together

Conditional probability is the mechanism by which information changes belief. The definition P(A|B) = P(A ∩ B) / P(B) restricts the sample space to B and renormalizes. The multiplication rule decomposes joint probabilities into chains of conditionals. Independence means conditioning carries no information. The law of total probability averages over a partition of the space. And Bayes’ theorem reverses the direction of conditioning, turning likelihoods into posteriors.

In the next lesson, we will apply these tools to real-world problems — medical testing, spam filtering, and decision-making under uncertainty — to see just how powerful and counterintuitive Bayesian reasoning can be.

Next up: Bayes’ Theorem in Practice — sensitivity, specificity, positive predictive value, the base-rate fallacy, and why even good tests can mislead when the disease is rare.

Key Takeaways
  • P(A|B) = P(A ∩ B) / P(B): conditioning on B restricts the sample space to B and renormalizes all probabilities.
  • The multiplication rule P(A ∩ B) = P(A|B) · P(B) decomposes joint probabilities into conditional chains.
  • Independence: P(A|B) = P(A), equivalently P(A ∩ B) = P(A) · P(B). Mutually exclusive events with positive probability are never independent.
  • The law of total probability computes P(A) by averaging P(A|Bi) over a partition of the sample space.
  • Bayes’ theorem inverts conditioning: P(A|B) = P(B|A) · P(A) / P(B).
  • P(A|B) ≠ P(B|A) in general — confusing them is the prosecutor’s fallacy and causes widespread errors in medical and legal reasoning.
Previous What is Probability? Module Overview Next Lesson Bayes’ Theorem in Practice