A Framework for Making Decisions Under Uncertainty
Science advances by testing ideas against evidence. You have a claim — a drug works, a coin is fair, a new algorithm is faster — and you gather data to evaluate it. But data are noisy. Even a perfectly fair coin can produce 7 heads in 10 flips. How do you distinguish a genuine effect from random chance?
Hypothesis testing provides a principled, replicable framework for making this judgment. Rather than asking “is the effect real?” directly (which requires knowing the true state of the world), it asks a more tractable question: how surprising would my data be if there were no effect? If the data would be very surprising under the “no effect” scenario, we have evidence against it.
This lesson covers the conceptual architecture of hypothesis testing: the null and alternative hypotheses, the test statistic, the p-value, the significance level, the two types of error, statistical power, and the choice between one-tailed and two-tailed tests.
The Null and Alternative Hypotheses
Every hypothesis test is built around two competing hypotheses. The null hypothesis, denoted H₀, represents the default, skeptical position — typically that there is no effect, no difference, or no relationship. The alternative hypothesis, denoted H₁ (or Ha), is what you are trying to find evidence for.
Drug trial: H₀: the drug has no effect on recovery time | H₁: the drug reduces recovery time
Coin fairness: H₀: p = 0.5 (coin is fair) | H₁: p ≠ 0.5 (coin is biased)
Algorithm speed: H₀: mean runtime = 100 ms | H₁: mean runtime < 100 ms
A critical asymmetry: we never prove H₀. We either reject H₀ (evidence against it is strong enough) or fail to reject H₀ (evidence is insufficient). This is analogous to a criminal trial — verdicts are “guilty” or “not guilty,” never “innocent.” Failing to reject H₀ does not mean H₀ is true; it means the data did not provide enough evidence to discredit it.
The Test Statistic
A test statistic is a single number computed from the data that summarizes the evidence against H₀. It is designed so that extreme values (far from what H₀ predicts) indicate strong evidence against H₀.
For testing a population mean μ when the population standard deviation σ is known, the natural test statistic is the Z-score of the sample mean:
Under H₀, this statistic follows a standard normal distribution (by the Central Limit Theorem). Large absolute values of Z indicate that the observed sample mean is far from what H₀ predicts — evidence against H₀.
Different testing scenarios call for different test statistics: the t-statistic when σ is unknown, the chi-squared statistic for categorical data, and the F-statistic for comparing multiple group means. The underlying logic is always the same: measure how far the data deviate from H₀.
The P-Value
The p-value is the probability, computed assuming H₀ is true, of observing a test statistic at least as extreme as the one actually observed. It quantifies how surprising the data are under the null hypothesis.
What the p-value is NOT: it is not the probability that H₀ is true, and it is not the probability that the result occurred by chance. These are common misinterpretations. The p-value is a property of the data given H₀, not a statement about H₀ given the data.
Wrong: “p = 0.03 means there is a 3% chance H₀ is true.”
Wrong: “p = 0.03 means there is a 97% chance the result is real.”
Correct: “If H₀ were true, I would see data this extreme or more extreme only 3% of the time.”
The Significance Level α
Before seeing the data, you choose a significance level α — the threshold below which a p-value is considered small enough to reject H₀. The most common choices are α = 0.05, α = 0.01, and α = 0.001.
The decision rule is simple: if p ≤ α, reject H₀ and declare the result “statistically significant.” If p > α, fail to reject H₀. The choice of α should be made before the experiment, not after seeing the data.
Why α = 0.05? There is nothing mathematically sacred about 0.05. It became the conventional threshold largely through historical precedent (R.A. Fisher used it in the 1920s). In fields where false discoveries are costly — particle physics, for instance — thresholds as small as 5 × 10−7 are used. In exploratory research, 0.1 is sometimes acceptable. The right threshold depends on the cost of a false positive versus the cost of a false negative.
Type I and Type II Errors
Any decision procedure in the presence of uncertainty will sometimes be wrong. Hypothesis testing makes two kinds of errors:
A Type I error (false positive) occurs when H₀ is actually true but we reject it. The probability of a Type I error is exactly α — by construction, we reject H₀ α of the time when it is true. This is the price of sensitivity.
A Type II error (false negative) occurs when H₀ is false but we fail to reject it. The probability of a Type II error is denoted β. Unlike α, β depends on the true effect size and sample size — it is not fixed by the significance threshold.
| H₀ is True | H₀ is False | |
|---|---|---|
| Reject H₀ | Type I Error (prob = α) | Correct Decision (prob = 1 − β) |
| Fail to Reject H₀ | Correct Decision (prob = 1 − α) | Type II Error (prob = β) |
There is an inherent trade-off: decreasing α (making it harder to reject H₀) reduces Type I errors but increases Type II errors. You cannot simultaneously minimize both types of error for a fixed sample size. The only way to reduce both is to increase the sample size.
Statistical Power
Power is the probability of correctly rejecting H₀ when it is false: Power = 1 − β. A test with high power is sensitive — it will reliably detect real effects. A test with low power frequently misses real effects, producing false negatives.
Power depends on four factors: the significance level α (higher α → more power), the true effect size (larger effects are easier to detect), the sample size n (more data → more power), and the variability in the data σ (less noise → more power). These relationships drive power analysis — the practice of computing the required sample size before collecting data to ensure the study can detect the effect of interest.
One-Tailed vs. Two-Tailed Tests
The alternative hypothesis determines the direction of the test. A two-tailed test considers deviations in both directions from H₀: H₁: μ ≠ μ0. The p-value is the probability of observing |Z| ≥ |zobs| in either tail of the distribution. This is the appropriate choice when you have no prior directional hypothesis.
A one-tailed test (directional test) considers only deviations in one direction: H₁: μ > μ0 (upper tail) or H₁: μ < μ0 (lower tail). The p-value is computed from only one tail, making it exactly half the two-tailed p-value for the same data. One-tailed tests have more power to detect the specified directional effect but provide no evidence about effects in the other direction.
Use two-tailed (default): you have no prior directional hypothesis, or you would care about effects in either direction (e.g., testing whether a new process changes yield at all).
Use one-tailed: the direction is theoretically motivated before data collection, and an effect in the opposite direction would be meaningless or uninteresting (e.g., testing whether a new drug reduces — not changes — blood pressure).
Warning: choosing a one-tailed test after seeing that the data favor one direction is p-hacking and invalidates the test.
The Complete Testing Procedure
A rigorous hypothesis test follows a fixed sequence of steps, all determined before data analysis:
1. State the hypotheses: Specify H₀ and H₁ in terms of the parameter of interest.
2. Choose the significance level: Select α before seeing the data (e.g., 0.05).
3. Select the test: Choose the appropriate test statistic for your data type and hypothesis.
4. Check assumptions: Verify that the conditions for the test (e.g., normality, independence) are met.
5. Compute the test statistic and p-value: Apply the formula to your data.
6. Decide and interpret: If p ≤ α, reject H₀. Report the effect size alongside the p-value.
- H₀ vs. H₁: H₀ is the skeptical default (no effect); H₁ is what you want to find evidence for. You never prove H₀ — you only reject it or fail to reject it.
- Test statistic: a number summarizing how far the data deviate from H₀. For a population mean test: Z = (x̄ − μ0) / (σ/√n).
- P-value: probability of data at least as extreme as observed, given H₀ is true. NOT the probability that H₀ is true.
- Significance level α: pre-chosen threshold (e.g., 0.05). Reject H₀ if p ≤ α.
- Type I error: rejecting a true H₀; probability = α. Type II error: failing to reject a false H₀; probability = β.
- Power = 1 − β: probability of detecting a real effect. Increases with larger n, larger effect size, and higher α.
- Two-tailed test: H₁: μ ≠ μ0 (default). One-tailed test: H₁: μ > or < μ0 (requires prior directional justification).