Conditional Probability
When we learn something about the world, probability changes. Conditional probability is the mathematics of “given that…”
What is P(A | B)?
Learning that B occurred restricts our world to outcomes inside B. The conditional probability of A given B is the fraction of B that also belongs to A.
The Formula
Divide the probability of both events occurring by the probability of the conditioning event. Requires P(B) > 0.
Multiplication Rule
Cross-multiply the definition to get the probability of both events occurring as a product of a conditional and a marginal.
Statistical Independence
A and B are independent when knowing B gives no information about A. The conditioning makes no difference.
Independent ≠ Disjoint
- Disjoint: A ∩ B = ∅ — they cannot both occur
- Independent: one occurring tells us nothing about the other
- If A ∩ B = ∅ and P(A), P(B) > 0 → they are dependent
- Knowing A happened means B definitely did not
Law of Total Probability
Partition Ω into mutually exclusive, exhaustive cases B1…Bn. Then P(A) is a weighted average of conditional probabilities.
Bayes’ Theorem
Turns a likelihood P(B|A) into a posterior P(A|B). Prior × likelihood ÷ evidence.
Tree Diagrams
- Each branch = one event, labeled with its probability
- Along a path: multiply probabilities
- Across paths: add probabilities
- Entire tree uses multiplication rule at every step
Total Probability at Work
Factory A: 60% of output, 2% defective. Factory B: 40%, 5% defective. Overall defect rate?
P(A|B) ≠ P(B|A)
- P(positive test | disease) = sensitivity (often high)
- P(disease | positive test) = depends on prevalence
- Prosecutor’s fallacy: P(evidence | innocent) ≠ P(innocent | evidence)
- Bayes’ theorem is the only safe way to invert conditioning
What you learned
Conditional probability, the multiplication rule, independence, the law of total probability, Bayes’ theorem, tree diagrams, and why P(A|B) ≠ P(B|A) — the engine of statistical inference.