Reading
Stories Mode

The Logic of Hypothesis Testing

~22 min read Lesson 1 of 4 in Module 6

A Framework for Making Decisions Under Uncertainty

Science advances by testing ideas against evidence. You have a claim — a drug works, a coin is fair, a new algorithm is faster — and you gather data to evaluate it. But data are noisy. Even a perfectly fair coin can produce 7 heads in 10 flips. How do you distinguish a genuine effect from random chance?

Hypothesis testing provides a principled, replicable framework for making this judgment. Rather than asking “is the effect real?” directly (which requires knowing the true state of the world), it asks a more tractable question: how surprising would my data be if there were no effect? If the data would be very surprising under the “no effect” scenario, we have evidence against it.

This lesson covers the conceptual architecture of hypothesis testing: the null and alternative hypotheses, the test statistic, the p-value, the significance level, the two types of error, statistical power, and the choice between one-tailed and two-tailed tests.

The Null and Alternative Hypotheses

Every hypothesis test is built around two competing hypotheses. The null hypothesis, denoted H₀, represents the default, skeptical position — typically that there is no effect, no difference, or no relationship. The alternative hypothesis, denoted H₁ (or Ha), is what you are trying to find evidence for.

Examples

Drug trial: H₀: the drug has no effect on recovery time  |  H₁: the drug reduces recovery time

Coin fairness: H₀: p = 0.5 (coin is fair)  |  H₁: p ≠ 0.5 (coin is biased)

Algorithm speed: H₀: mean runtime = 100 ms  |  H₁: mean runtime < 100 ms

A critical asymmetry: we never prove H₀. We either reject H₀ (evidence against it is strong enough) or fail to reject H₀ (evidence is insufficient). This is analogous to a criminal trial — verdicts are “guilty” or “not guilty,” never “innocent.” Failing to reject H₀ does not mean H₀ is true; it means the data did not provide enough evidence to discredit it.

The Test Statistic

A test statistic is a single number computed from the data that summarizes the evidence against H₀. It is designed so that extreme values (far from what H₀ predicts) indicate strong evidence against H₀.

For testing a population mean μ when the population standard deviation σ is known, the natural test statistic is the Z-score of the sample mean:

Z Test Statistic
Z = \dfrac{\bar{x} - \mu_0}{\sigma / \sqrt{n}}
x̄ is the sample mean, μ₀ is the hypothesized mean under H₀, σ is the known population standard deviation, and n is the sample size. Z measures how many standard errors the sample mean is from the hypothesized value.

Under H₀, this statistic follows a standard normal distribution (by the Central Limit Theorem). Large absolute values of Z indicate that the observed sample mean is far from what H₀ predicts — evidence against H₀.

Different testing scenarios call for different test statistics: the t-statistic when σ is unknown, the chi-squared statistic for categorical data, and the F-statistic for comparing multiple group means. The underlying logic is always the same: measure how far the data deviate from H₀.

The P-Value

The p-value is the probability, computed assuming H₀ is true, of observing a test statistic at least as extreme as the one actually observed. It quantifies how surprising the data are under the null hypothesis.

P-Value (Two-Tailed)
p\text{-value} = P\bigl(|Z| \geq |z_{\text{obs}}| \;\big|\; H_0\bigr)
For a two-tailed test, the p-value is the probability of observing |Z| ≥ |z_obs| under H₀. A small p-value means the observed data would be rare if H₀ were true.

What the p-value is NOT: it is not the probability that H₀ is true, and it is not the probability that the result occurred by chance. These are common misinterpretations. The p-value is a property of the data given H₀, not a statement about H₀ given the data.

Common Misconceptions

Wrong: “p = 0.03 means there is a 3% chance H₀ is true.”

Wrong: “p = 0.03 means there is a 97% chance the result is real.”

Correct: “If H₀ were true, I would see data this extreme or more extreme only 3% of the time.”

The Significance Level α

Before seeing the data, you choose a significance level α — the threshold below which a p-value is considered small enough to reject H₀. The most common choices are α = 0.05, α = 0.01, and α = 0.001.

The decision rule is simple: if p ≤ α, reject H₀ and declare the result “statistically significant.” If p > α, fail to reject H₀. The choice of α should be made before the experiment, not after seeing the data.

Why α = 0.05? There is nothing mathematically sacred about 0.05. It became the conventional threshold largely through historical precedent (R.A. Fisher used it in the 1920s). In fields where false discoveries are costly — particle physics, for instance — thresholds as small as 5 × 10−7 are used. In exploratory research, 0.1 is sometimes acceptable. The right threshold depends on the cost of a false positive versus the cost of a false negative.

Type I and Type II Errors

Any decision procedure in the presence of uncertainty will sometimes be wrong. Hypothesis testing makes two kinds of errors:

A Type I error (false positive) occurs when H₀ is actually true but we reject it. The probability of a Type I error is exactly α — by construction, we reject H₀ α of the time when it is true. This is the price of sensitivity.

A Type II error (false negative) occurs when H₀ is false but we fail to reject it. The probability of a Type II error is denoted β. Unlike α, β depends on the true effect size and sample size — it is not fixed by the significance threshold.

H₀ is True H₀ is False
Reject H₀ Type I Error (prob = α) Correct Decision (prob = 1 − β)
Fail to Reject H₀ Correct Decision (prob = 1 − α) Type II Error (prob = β)

There is an inherent trade-off: decreasing α (making it harder to reject H₀) reduces Type I errors but increases Type II errors. You cannot simultaneously minimize both types of error for a fixed sample size. The only way to reduce both is to increase the sample size.

Statistical Power

Power is the probability of correctly rejecting H₀ when it is false: Power = 1 − β. A test with high power is sensitive — it will reliably detect real effects. A test with low power frequently misses real effects, producing false negatives.

Statistical Power
\text{Power} = 1 - \beta = P(\text{reject } H_0 \mid H_0 \text{ is false})
Power is the complement of the Type II error rate. A conventional target is 80% power (β = 0.20), though 90% is preferred in high-stakes research.

Power depends on four factors: the significance level α (higher α → more power), the true effect size (larger effects are easier to detect), the sample size n (more data → more power), and the variability in the data σ (less noise → more power). These relationships drive power analysis — the practice of computing the required sample size before collecting data to ensure the study can detect the effect of interest.

One-Tailed vs. Two-Tailed Tests

The alternative hypothesis determines the direction of the test. A two-tailed test considers deviations in both directions from H₀: H₁: μ ≠ μ0. The p-value is the probability of observing |Z| ≥ |zobs| in either tail of the distribution. This is the appropriate choice when you have no prior directional hypothesis.

A one-tailed test (directional test) considers only deviations in one direction: H₁: μ > μ0 (upper tail) or H₁: μ < μ0 (lower tail). The p-value is computed from only one tail, making it exactly half the two-tailed p-value for the same data. One-tailed tests have more power to detect the specified directional effect but provide no evidence about effects in the other direction.

When to Use One-Tailed Tests

Use two-tailed (default): you have no prior directional hypothesis, or you would care about effects in either direction (e.g., testing whether a new process changes yield at all).

Use one-tailed: the direction is theoretically motivated before data collection, and an effect in the opposite direction would be meaningless or uninteresting (e.g., testing whether a new drug reduces — not changes — blood pressure).

Warning: choosing a one-tailed test after seeing that the data favor one direction is p-hacking and invalidates the test.

The Complete Testing Procedure

A rigorous hypothesis test follows a fixed sequence of steps, all determined before data analysis:

The Six Steps

1. State the hypotheses: Specify H₀ and H₁ in terms of the parameter of interest.

2. Choose the significance level: Select α before seeing the data (e.g., 0.05).

3. Select the test: Choose the appropriate test statistic for your data type and hypothesis.

4. Check assumptions: Verify that the conditions for the test (e.g., normality, independence) are met.

5. Compute the test statistic and p-value: Apply the formula to your data.

6. Decide and interpret: If p ≤ α, reject H₀. Report the effect size alongside the p-value.

Key Takeaways
  • H₀ vs. H₁: H₀ is the skeptical default (no effect); H₁ is what you want to find evidence for. You never prove H₀ — you only reject it or fail to reject it.
  • Test statistic: a number summarizing how far the data deviate from H₀. For a population mean test: Z = (x̄ − μ0) / (σ/√n).
  • P-value: probability of data at least as extreme as observed, given H₀ is true. NOT the probability that H₀ is true.
  • Significance level α: pre-chosen threshold (e.g., 0.05). Reject H₀ if p ≤ α.
  • Type I error: rejecting a true H₀; probability = α. Type II error: failing to reject a false H₀; probability = β.
  • Power = 1 − β: probability of detecting a real effect. Increases with larger n, larger effect size, and higher α.
  • Two-tailed test: H₁: μ ≠ μ0 (default). One-tailed test: H₁: μ > or < μ0 (requires prior directional justification).
Previous Bootstrap Methods Module Overview Next Common Tests