The Logic of Hypothesis Testing
How do you tell a real effect from random noise? You ask how surprised you’d be if there were no effect at all. That’s the entire logic of hypothesis testing.
H₀ vs. H₁
H₀ (null): the skeptical default — no effect, no difference. H₁ (alternative): what you are trying to find evidence for.
The Test Statistic
Counts how many standard errors the sample mean is from the hypothesized value. Large |Z| = strong evidence against H₀.
The P-Value
Probability of observing data this extreme (or more), assuming H₀ is true.
Significance Level α
α = 0.01
α = 0.001
p > α → fail to reject
Type I Error
Rejecting H₀ when it is actually true. The probability is exactly α — you set this yourself.
Type II Error
Failing to reject H₀ when it is actually false. Probability = β. Unlike α, β is not fixed — it depends on the effect size and sample size.
Larger true effect
Less noise
Statistical Power
Power = 1 − β. The probability of correctly rejecting a false H₀. Conventional target: 80% (sometimes 90% in high-stakes research).
- Larger sample size → more power
- Larger effect size → easier to detect
- Higher α → more power (but more false positives)
- Less variance → sharper signal
One-Tailed vs. Two-Tailed
Six Steps
- State H₀ and H₁
- Choose α before seeing data
- Select the appropriate test
- Check assumptions
- Compute test statistic and p-value
- Decide: reject or fail to reject. Report effect size too.
What you learned
- H₀: no effect (default). H₁: the claim you’re testing
- Test statistic measures deviation from H₀
- P-value = P(data this extreme | H₀ true) — not P(H₀ is true)
- Reject H₀ when p ≤ α (set in advance)
- Type I error (α): false positive. Type II error (β): false negative
- Power = 1 − β: target ≥ 80%