LA 101
M12 · Project
Capstone project
Reduce first, then model

Chain two sandboxes into one data-analysis pipeline — first collapse a wide dataset to the few principal components that carry its variance, read how many components and the reduction ratio that buys; then fit an ordinary least-squares model and read how much of its variance the fit explains. Predict each number before the screen shows it, and record your pipeline at the end.

Two stages, one quantity: variance explained
\begin{gathered} k^\ast = \min\{k : \mathrm{EVR}(k) \ge 0.95\} \quad\text{(reduce)} \\[4pt] R^2 = 1 - \mathrm{SSE}/\mathrm{SST} \quad\text{(model)} \end{gathered}

A real analysis is a pipeline, and both ends of it speak the same language: fraction of variance explained. Principal-component analysis rotates the axes to the directions the data varies in most, so the explained-variance ratio EVR(k) says how much variance the top k components keep — and you drop the rest. The model that follows is judged by the very same idea: R² = 1 - SSE/SST is the fraction of the response's variance the fit explains. This project is the interaction: raise the variance threshold and k* grows; and a fit that explains most of the variance is exactly a high R². Reduction and regression are two views of one quantity.

1 / 8
LA 101
M12 · Project
Set it up
Two sandboxes, one pipeline

Open both sandboxes in tabs. Each step below names the one sandbox to read and the single choice to make in it; leave every other knob at its default so your numbers match these. The reduction runs on the “Sharp geometry” image; the model runs on the “Clean linear” preset.

The three stages
1 PCA / SVD (LA 101) -> Image = Sharp geometry, 64x64, seed 1, threshold 95% 2 Regression (Stat 101) -> Preset = Clean linear, seed 3, fit OLS y = b0 + b1 x

Read four numbers across the two tabs: the smallest number of components reaching 95% variance and the reduction ratio it buys in the PCA explorer, then the coefficient of determination and the fitted slope in the regression workshop.

2 / 8
LA 101
M12 · Project
Step 1 of 4
Demo · PCA / SVD explorer
Reduce: how many PCs for 95%

In the PCA / SVD explorer choose the “Sharp geometry” image (defaults: 64×64, no noise, seed 1). This image has 40 non-zero components. Step k up and watch “energy kept” cross 95%; predict the smallest k that reaches it.

Expected

Just 3 principal components reach 95% of the variance (EVR = 0.9549) — so the pipeline keeps 3 directions and drops the other 61. This is the demo's own svd() / explained-variance arithmetic, not a typed number.

3 / 8
LA 101
M12 · Project
Step 2 of 4
Demo · PCA / SVD explorer
Reduce: the reduction ratio

Keeping k* = 3 of the 64 dimensions is a reduction ratio n / k* = 64 / 3. Predict that ratio — how many times smaller the representation is while it still holds 95% of the variance.

Expected

The ratio is 64 / 3 = 21.33× — three principal components stand in for 64 original dimensions. That smaller, denser representation is what feeds the model in stage two. This is the demo's own svd() / explained-variance arithmetic.

4 / 8
LA 101
M12 · Project
Step 3 of 4
Demo · Regression workshop
Model: how much variance the fit explains

Switch to the regression workshop (a Statistics 101 sandbox). Load the “Clean linear” preset with seed 3 and fit ordinary least squares. The coefficient of determination R² = 1 - SSE/SST is the fraction of the response's variance the line explains — the same “variance explained” idea as stage one. Predict whether it is close to 1.

Expected

The fit explains R² = 0.9429 of the variance — about 94%, a tight linear relationship. Read against stage one, the parallel is exact: PCA kept 95% of the variance in three directions, and the model captures 94% of the response's. This is the demo's own analyseFit() output, not a typed number.

5 / 8
LA 101
M12 · Project
Step 4 of 4
Demo · Regression workshop
Model: the fitted slope

Same “Clean linear” preset, seed 3. Beyond how well the line fits, its slope b1 is the effect size — how much y moves per unit of x. Predict roughly the slope of the fitted line, then read b1.

Expected

The ordinary-least-squares slope is b1 = 1.2508 — each unit of x lifts y by about 1.25. Slope and R² are different facts: the slope is the effect, the R² is how tightly the points hug it. This is the demo's own analyseFit() arithmetic.

6 / 8
LA 101
M12 · Project
Your turn
Assemble the pipeline

Open both sandboxes and run the pipeline yourself: raise the variance threshold and watch how many principal components you must keep grow, then fit the model and watch its R-squared measure the same variance-explained idea on the response side.

7 / 8
LA 101
M12 · Project
Deliverable
Record your pipeline

A pipeline is a set of stages and the numbers they produce. Write yours down — fill in each blank from the sandboxes, then add one sentence of reasoning. This is the deliverable; there is no number to read off the screen here.

Write this down
Stage 1 PCA / SVD ...... Sharp geometry, 95% variance k* components .. ______ reduction n/k* . ______ x Stage 2 Regression ..... Clean linear, seed 3, OLS R-squared ...... ______ slope b1 ....... ______ Pipeline note ........... how does the model's R-squared echo the PCA variance? _____ Design note ............. which knob would you change first, and why? __________

Then change one thing — a higher variance threshold, a noisier regression preset — and note which of the four numbers moved and by how much. That coupling, and the shared “variance explained” language of both stages, is the whole lesson of the project.

8 / 8