Reading
Stories Mode

Simple Linear Regression

~28 min read Lesson 1 of 4 in Module 7

From Scatter to Straight Line

Data rarely arranges itself perfectly. Observations scatter across a plot, driven partly by the relationship we care about and partly by noise we cannot control. Simple linear regression is the tool that cuts through that scatter to reveal the underlying linear trend — and to quantify how confidently we can use one variable to predict another.

The word “simple” here means one predictor. We have a single input variable X (also called the independent variable, predictor, or feature) and a single output Y (the dependent variable or response). The goal is to find the straight line that best describes their relationship and to make that notion of “best” precise.

The Statistical Model

Simple linear regression models Y as a linear function of X plus irreducible random noise:

Simple Linear Regression Model
Y = \beta_0 + \beta_1 X + \varepsilon
Y is the response, β₀ the intercept, β₁ the slope, X the predictor, and ε the error term — assumed to be independent with mean zero and constant variance σ².

The parameters β₀ and β₁ are fixed but unknown constants that describe the true population relationship. The error term ε captures everything that X does not explain: measurement error, omitted variables, inherent randomness. Because ε is random, each observation of Y is also random, even for a fixed X.

The model makes four key assumptions: (1) the relationship between X and Y is linear; (2) the errors are independent of each other; (3) the errors have mean zero; (4) the errors have constant variance σ² (homoscedasticity). We will return to each assumption when we discuss residual analysis.

Ordinary Least Squares Estimation

We observe n data pairs (x₁, y₁), …, (xₙ, yₙ). From these we want to estimate β₀ and β₁. The Ordinary Least Squares (OLS) principle selects the estimates β̂₀ and β̂₁ that minimize the total squared vertical distance between the observed points and the fitted line:

OLS Objective
\text{SSR} = \sum_{i=1}^{n}(y_i - \hat{y}_i)^2
SSR = Sum of Squared Residuals. Each residual eᵢ = yᵢ − ŷᵢ is the vertical gap between the observed value and the value predicted by the line at xᵢ.

Taking derivatives of SSR with respect to β₀ and β₁ and setting them to zero yields the normal equations, whose solution is:

OLS Slope Estimator
\hat{\beta}_1 = \dfrac{\sum_{i=1}^{n}(x_i-\bar{x})(y_i-\bar{y})}{\sum_{i=1}^{n}(x_i-\bar{x})^2}
β̂₁ equals the sample covariance of X and Y divided by the sample variance of X. The numerator measures how X and Y move together; the denominator scales by the spread of X.
OLS Intercept Estimator
\hat{\beta}_0 = \bar{y} - \hat{\beta}_1\bar{x}
The fitted line always passes through the point (x̄, ȳ) — the center of mass of the data. The intercept is determined by this constraint and the estimated slope.

Under the four model assumptions, OLS estimators are BLUE — Best Linear Unbiased Estimators (Gauss-Markov theorem). They are unbiased (on average equal to the true parameters) and have the smallest variance among all linear unbiased estimators.

Interpreting the Coefficients

Careful interpretation of β̂₀ and β̂₁ is essential before drawing any conclusions from a regression.

The slope β̂₁ is the most important quantity. It tells us: for a one-unit increase in X, the predicted value of Y changes by β̂₁ units, on average, holding all else equal. The sign indicates direction (positive or negative association); the magnitude indicates strength. Note that correlation drives direction but not causation — regression coefficients describe statistical association, not causal mechanisms, unless the data come from a randomized experiment.

The intercept β̂₀ is the predicted Y when X = 0. This is meaningful only if X = 0 is a plausible value within the observed range of the data. If you are regressing salary on years of experience and your data range from 1 to 30 years, the intercept (salary at zero years) is an extrapolation and should be interpreted cautiously or not at all.

Worked Example

Suppose we regress house price (in thousands of dollars) on floor area (in square feet) and obtain: Pricê = 48.3 + 0.124 × Area. The slope says each additional square foot adds $124 to the predicted price. The intercept ($48,300) is the predicted price for a zero-square-foot house — a meaningless extrapolation, but mathematically necessary to anchor the line. The useful takeaway is the slope.

Measuring Fit: R²

Once we have a fitted line, we need to measure how well it describes the data. The coefficient of determination R² does this by comparing the variation explained by the model to the total variation in Y:

R-squared
R^2 = 1 - \dfrac{\text{SSR}}{\text{SST}} = 1 - \dfrac{\displaystyle\sum_{i=1}^{n}(y_i-\hat{y}_i)^2}{\displaystyle\sum_{i=1}^{n}(y_i-\bar{y})^2}
SST = total sum of squares = Σ(yᵢ − ȳ)². SSR = residual sum of squares = Σ(yᵢ − ŷᵢ)². R² ranges from 0 (line explains nothing) to 1 (perfect fit). In simple linear regression, R² equals the square of the Pearson correlation coefficient r.

R² = 0.75 means the regression line accounts for 75% of the variability in Y; the remaining 25% is unexplained by X. High R² indicates a good fit, but “good” is domain-specific: R² = 0.3 might be excellent for predicting individual stock returns and terrible for predicting a manufacturing yield controlled to tight tolerances.

R² Does Not Prove Linearity

Anscombe’s Quartet is a famous collection of four datasets that all share nearly identical R², slope, and intercept — yet look completely different when plotted. One is perfectly linear, one is a perfect parabola fit by a line, one has a single outlier distorting an otherwise perfect fit, and one has no real relationship at all. Always plot your data. R² alone cannot distinguish these situations.

Residual Analysis: Checking Assumptions

The validity of OLS inference depends on the four model assumptions. Residual analysis — examining the pattern in eᵢ = yᵢ − ŷᵢ — is the primary diagnostic tool.

Linearity: Plot residuals against the fitted values ŷ or against X. If the relationship is truly linear, residuals should scatter randomly around zero with no systematic curve or pattern. A U-shaped or inverted-U pattern indicates the true relationship is nonlinear and a linear model is misspecified.

Normality of residuals: Many inferential procedures (confidence intervals, t-tests on coefficients) rely on residuals being approximately normal. Check with a histogram of residuals or, better, a QQ plot. Points falling along the diagonal suggest normality; systematic deviations suggest skewness or heavy tails. Note: with large n, the Central Limit Theorem provides some protection, but extreme non-normality can still distort inference in small samples.

Homoscedasticity (constant variance): The spread of residuals should be roughly the same across all values of X. On a residual-vs-fitted plot, look for a consistent vertical spread. A “fanning out” pattern — where residuals grow larger as ŷ increases — signals heteroscedasticity. This does not bias the coefficient estimates but makes standard errors unreliable, invalidating hypothesis tests and confidence intervals.

Homoscedasticity Assumption
\text{Var}(\varepsilon_i \mid X = x) = \sigma^2 \quad \forall\, x
The variance of ε is σ² regardless of the value of X. Violation (heteroscedasticity) means OLS standard errors are incorrect even though coefficient estimates remain unbiased.

Independence: Residuals should be uncorrelated with each other. Independence is violated most often in time-series data, where consecutive observations tend to be similar (autocorrelation). Durbin-Watson test or a plot of residuals against observation order can reveal this. Correlated errors understate uncertainty and produce overly narrow confidence intervals.

Confidence Intervals vs. Prediction Intervals

A fitted line ŷ = β̂₀ + β̂₁x gives a point estimate for a given x, but we also need an interval to capture uncertainty. Two distinct intervals serve different purposes:

A confidence interval for the mean response at x = x* answers: “Where does the true average value of Y lie when X = x*?” It captures uncertainty from estimating β₀ and β₁ from data, but not the irreducible noise ε.

A prediction interval for a new observation at x = x* answers: “If I observe a new data point at X = x*, where will it land?” It must account for both the uncertainty in estimating the mean response and the additional scatter due to ε. Prediction intervals are always wider than confidence intervals, and they do not shrink to zero even with infinite data, because ε is irreducible.

Both Intervals Widen Away from x̄

Both the confidence and prediction intervals are narrowest at x = x̄ (the mean of the observed X values) and widen as x* moves away from x̄. This reflects the fact that extrapolation increases uncertainty: estimates of the regression line become less reliable the further you venture from the center of the data. Never extrapolate blindly.

Testing the Slope: Is X a Useful Predictor?

The most important inferential question in simple linear regression is whether the slope is significantly different from zero. If β₁ = 0, then X provides no linear information about Y and the regression is useless as a predictive model.

We test H₀: β₁ = 0 against H₁: β₁ ≠ 0 using a t-statistic:

t-test for Slope
t = \dfrac{\hat{\beta}_1}{\text{SE}(\hat{\beta}_1)} \sim t_{n-2} \text{ under } H_0
Under H₀, t follows a t-distribution with n−2 degrees of freedom. SE(β̂₁) = s/√Σ(xᵢ−x̄)², where s = √(SSR/(n−2)) is the residual standard error.

Reject H₀ if |t| > t₋₍ᴰ₎ₓ(n − 2) at significance level α. The equivalent F-test (which generalizes naturally to multiple regression) yields F = t² and produces the same p-value. A significant result means X is a statistically useful linear predictor of Y — not that the relationship is large, causal, or practically important. Always pair the test with R² and a confidence interval for β₁.

The 95% confidence interval for β₁ is β̂₁ ± t₋₍ᴰ₎ₓ(n − 2) × SE(β̂₁). An interval that does not include zero is equivalent to rejecting H₀ at α = 0.05. The interval also quantifies the plausible magnitude of the effect, which is often more informative than the binary reject/fail-to-reject decision.

Regression vs. Correlation

The Pearson correlation r and the regression slope β̂₁ are related by β̂₁ = r × (sʏ/sₓ), where sʏ and sₓ are the sample standard deviations of Y and X. They test the same null hypothesis H₀: ρ = 0 and produce the same p-value. The difference is interpretive: correlation is symmetric (swapping X and Y leaves r unchanged), while regression is directional (Y depends on X) and produces a line for prediction.

Key Takeaways
  • The model Y = β₀ + β₁X + ε describes Y as a linear function of X plus random error with mean zero and constant variance.
  • OLS minimizes the sum of squared residuals to produce unbiased, minimum-variance estimators β̂₀ and β̂₁ (Gauss-Markov theorem).
  • The slope β̂₁ is the average change in Y per unit increase in X. The intercept is meaningful only if X = 0 is within the data range.
  • R² = 1 − SSR/SST measures the proportion of Y’s variance explained by the model. It ranges from 0 to 1 but must be interpreted in domain context.
  • Residual plots diagnose assumption violations: random scatter is good; patterns indicate nonlinearity, heteroscedasticity, or autocorrelation.
  • Prediction intervals are always wider than confidence intervals at the same x* because they include irreducible error ε that does not vanish even with infinite data.
  • Testing H₀: β₁ = 0 via the t-statistic determines whether X is a statistically useful linear predictor. Always report the effect size and a CI alongside the p-value.
Previous Common Pitfalls Module Overview Next Multiple Regression