AI-generated Computed, not drawn Thresholds are conventions

This diagnostics workshop was generated by Claude (Anthropic) for the STAT 101 course materials. Every number on it is computed in your browser and recomputed on every drag: the least-squares coefficients, all residuals, s² with the n − P divisor, the standard errors, the t statistic and its two-sided p-value from a real incomplete-beta evaluation, R² two independent ways and cross-checked, every hat value, every studentised residual, the confidence and prediction bands, and the QQ plot's normal scores from a real inverse normal CDF.

Cook's distance is computed twice on purpose. Once by the algebraic shortcut using leverage and the residual, and once by its definition — dropping the point, refitting, and adding up how far every fitted value moved. Comparison 3 prints the difference, so the shortcut is demonstrated rather than asserted. This matters here: the denominator must be the number of parameters (P = 2), and using the number of predictors instead makes it exactly twice too large. The lessons had it the second way and were corrected on all four surfaces on 2026-08-23 (see TODO.md §10). The page prints Σhii = P live so you can see where that 2 comes from.

The three flags are conventions, not facts, and are labelled as such at the point of use: leverage > 2P/n ("twice the average"; some texts use 3P/n), Cook's D > 1 (Cook 1977, and generous in practice), and |studentised residual| > 2 (a rough normal cut-off). Every actual number is printed beside its flag so you can disagree with the threshold. Nothing on this page rests on a figure transcribed from a document.

Chosen rather than computed: the six datasets and the noise seed — though the noise is a real seeded Gaussian draw, so its statistics are computed and every preset is reproducible. Known limits, measured and stated rather than hidden: the inverse normal used for the QQ plot is accurate to about 1.8×10−9, and the normal CDF to about 1.2×10−7. A Newton-style refinement of the inverse was written and then removed, because measuring it against 15-digit reference quantiles showed it made the answer thirty times worse — it was correcting towards the coarser of the two functions. The checker asserts the accuracy actually achieved. Scope: simple linear regression only — the multiple-regression builder and the logistic half of the syllabus lab are not here.

Regression influence — leverage, Cook's distance, and what actually moves the line

Course demo — linked from the Module 7 lesson decks in both languages; the page itself is English‑only for now. Built for STAT-101 Lessons 7.1 and 7.3 and for the scatter-plus-diagnostics half of the lab those lessons declare. Module 7 has four decks and no canvases at all, so the residual plot, the QQ plot, heteroscedasticity, leverage and Cook's distance — five things that are entirely visual — are described and never shown. Drag the points. The one thing to take away: leverage is opportunity, influence is opportunity taken, and the two presets that differ only in one y value prove they are not the same thing.

 

 

 

 

 
 

 

 
 

 

 
 

 

 
 

  

 

 

3

 

  

11

 

 

  

 

 

 — 
 — 
 — 
 — 
 — 
 — 
 — 
 — 

1 

 

 

 

2 

 

 

3 

 

 

4 

 

 

5 

 

 

The arithmetic, in full

The chain in words