Reading
Stories Mode

Designing Experiments

~13 min read Lesson 1 of 3 in Module 11

From Observation to Intervention

Statistics is often applied to data that already exists — surveys, records, observational studies. But the most powerful way to answer a causal question is to design an experiment: deliberately intervene in the world, control what you can, and measure the result. Well-designed experiments allow us to make causal claims that observational data cannot.

The difference is fundamental. An observational study might find that people who drink coffee have lower rates of depression. But coffee drinkers may also exercise more, sleep better, or have higher incomes. Any of these could explain the correlation. Only an experiment — randomly assigning some people to drink coffee and others not — can isolate the causal effect of coffee itself.

The Core Goal

Experimental design is the art of arranging a study so that differences in outcomes can be attributed to the treatment, and not to anything else. Every design principle — randomization, control groups, blocking — exists to guard against this single threat: confounding.

Confounding = a third variable that is associated with both the treatment and the outcome

Randomization: The Cornerstone

Randomization is the single most important tool in experimental design. When participants are randomly assigned to treatment and control groups, all confounding variables — both measured and unmeasured — are balanced across groups on average. This is the magic of randomization: it controls for everything you thought of and everything you did not think of.

Without randomization, a clinical trial might assign healthier patients to the treatment group (consciously or unconsciously), making the treatment look more effective than it is. With randomization, the two groups are expected to be similar in age, health status, lifestyle, and every other factor. Any difference in outcomes can then be attributed to the treatment alone.

Randomization can be implemented in several ways: simple random assignment (like flipping a coin), stratified randomization (ensuring balance within subgroups), or cluster randomization (randomizing groups rather than individuals). The choice depends on the structure of the study and what confounds are most worrying.

Control Groups

Every experiment needs a control group: a group that does not receive the treatment, providing a baseline for comparison. Without a control group, we cannot know whether any change we observe is due to the treatment or would have happened anyway (natural recovery, seasonal variation, regression to the mean).

In medicine, the control group often receives a placebo — an inert treatment that looks identical to the real one. This is essential because patients who believe they are being treated often improve even when the treatment is chemically inactive. The placebo effect is real and measurable; blinding participants to their treatment assignment controls for it.

A double-blind design goes further: neither the participants nor the researchers administering the treatment know who is in which group. This prevents both the placebo effect (on participants) and experimenter bias (on the researchers who might unconsciously treat or measure the two groups differently).

Blocking: Reducing Variability

Blocking is a technique for increasing the precision of an experiment by controlling for a known source of variability. A block is a group of experimental units that are similar to each other with respect to some characteristic that is expected to affect the outcome.

Suppose you are testing a new fertilizer and you have plots of land on two different soil types. Soil type will affect crop yield regardless of the fertilizer. If you ignore soil type, it adds noise to your estimate of the fertilizer effect. Instead, you can block by soil type: within each soil type, randomly assign plots to treatment and control. Then compare treatment vs. control within each block, and combine the within-block comparisons. The soil-type variability is removed, and your estimate of the fertilizer effect is more precise.

The Blocking Principle

Block what you can; randomize what you cannot.

Known confounders should be controlled by blocking. Unknown confounders are handled by randomization. Together, they eliminate both the bias (from known confounders) and average out the bias (from unknown ones).

The most extreme form of blocking is a matched pairs design, where each unit is paired with the most similar unit in the other treatment group — or where the same unit receives both treatments in sequence (a crossover design). This removes all between-subject variability from the comparison.

Factorial Designs

A factorial design studies two or more factors simultaneously in a single experiment. Instead of running separate experiments for each factor, you cross all levels of every factor with all levels of every other factor. This is far more efficient and, crucially, allows you to detect interaction effects — cases where the effect of one factor depends on the level of another.

Consider testing a new drug at two doses (low, high) and in combination with two diets (standard, low-sodium). A full 2×2 factorial design would have four groups: (low dose, standard), (low dose, low-sodium), (high dose, standard), (high dose, low-sodium). This design tells you the main effect of dose, the main effect of diet, and whether the drug works better with a particular diet — all from a single experiment.

When many factors are involved, a fractional factorial design tests only a carefully chosen subset of all combinations, sacrificing the ability to estimate some high-order interactions in exchange for dramatically fewer experimental units. This is common in industrial process optimization and early-phase drug development.

Sample Size and Power Analysis

How many participants do you need? This question has a precise statistical answer. Power analysis determines the minimum sample size required to detect an effect of a given size with a given probability (the power) at a given significance level.

Power is the probability of correctly rejecting a false null hypothesis — of detecting a real effect when one exists. A study with low power is likely to miss a real effect, producing a false negative. Conventionally, power is set at 80% or 90%, meaning a 20% or 10% chance of missing a true effect.

Sample Size Formula (Two-Group Comparison)
n = \frac{2\,\sigma^2\,(z_{\alpha/2} + z_{\beta})^2}{\delta^2}
n is the required sample size per group; zα/2 is the critical value for significance level α; zβ is the critical value for power (1−β); σ² is the variance; δ is the minimum detectable effect size.

Four quantities are linked: sample size, effect size, significance level, and power. Fixing any three determines the fourth. A smaller effect size requires a larger sample to detect. A more stringent significance level (lower α) requires more data. Higher desired power requires more data. In practice, you specify the effect size you care about (the smallest difference that would be practically meaningful), choose α = 0.05 and power = 0.80, and solve for n.

Effect Size (Cohen's d)
d = \frac{\mu_1 - \mu_2}{\sigma}
Standardized difference between means. Small: 0.2, medium: 0.5, large: 0.8.
Power
\text{Power} = 1 - \beta = P(\text{reject } H_0 \mid H_1 \text{ true})
Probability of detecting a true effect. 1 − β where β is the Type II error rate.

Underpowered studies waste resources, produce unreliable estimates, and contribute to the replication crisis. Pre-registering your power analysis — specifying your sample size calculation before collecting data — prevents post-hoc adjustments and increases trust in your results.

Ethical Considerations

Experiments involve deliberately giving some people a treatment and withholding it from others. This raises ethical obligations that must be addressed before any experiment begins.

Equipoise is the principle that an experiment is only ethical when there is genuine uncertainty about which treatment is better. If you already know one treatment is superior, it is unethical to randomly assign patients to the inferior one. Experiments should stop early (via pre-specified stopping rules) if evidence accumulates strongly enough in favor of one arm.

Informed consent requires that participants understand what the experiment involves, what the risks are, and that they are free to withdraw at any time without penalty. Deception is sometimes necessary (e.g., blinding participants to treatment assignment), but it must be disclosed after the study, and its use must be justified.

Institutional Review Boards (IRBs) — or ethics committees — review and approve experimental designs involving human subjects. They assess risk-benefit trade-offs, require appropriate consent procedures, and ensure vulnerable populations are protected. All human-subjects research must obtain IRB approval before data collection begins.

Good Experimental Practice Checklist

1. Pre-register — specify hypothesis, sample size, and analysis plan before collecting data.

2. Randomize — use a verifiable random process for treatment assignment.

3. Blind — mask treatment assignment from participants and, where possible, from researchers.

4. Block — control for known confounders by grouping similar units.

5. Power — compute required sample size before starting; don't stop early without a pre-specified rule.

6. Report completely — publish regardless of outcome; report all measured variables.

In the next lesson, we will apply these principles to the specific context of A/B testing — the experimental framework used by technology companies to evaluate product changes at massive scale. You will see how randomization, sample size, and stopping rules play out in a digital setting where thousands of experiments run simultaneously.

Key Takeaways
  • Experiments allow causal inference; observational studies cannot rule out confounding. Randomization is the key that unlocks causation.
  • Control groups (and placebo controls in medicine) provide the baseline against which treatment effects are measured. Blinding prevents placebo effects and experimenter bias.
  • Blocking controls for known sources of variability, increasing the precision of treatment estimates without requiring more subjects.
  • Factorial designs study multiple factors simultaneously and reveal interaction effects that separate experiments would miss.
  • Power analysis determines the required sample size before data collection; underpowered studies are a primary driver of irreproducible findings.
  • Ethical constraints — equipoise, informed consent, IRB approval — are not obstacles but essential safeguards for responsible science.
Previous Forecasting Module Overview Next Lesson A/B Testing