Chain two sandboxes into one data-analysis pipeline — first collapse a wide dataset to the few principal components that carry its variance, read how many components and the reduction ratio that buys; then fit an ordinary least-squares model and read how much of its variance the fit explains. Predict each number before the screen shows it, and record your pipeline at the end.
A real analysis is a pipeline, and both ends of it speak the same language: fraction of variance explained. Principal-component analysis rotates the axes to the directions the data varies in most, so the explained-variance ratio EVR(k) says how much variance the top k components keep — and you drop the rest. The model that follows is judged by the very same idea: R² = 1 - SSE/SST is the fraction of the response's variance the fit explains. This project is the interaction: raise the variance threshold and k* grows; and a fit that explains most of the variance is exactly a high R². Reduction and regression are two views of one quantity.
Open both sandboxes in tabs. Each step below names the one sandbox to read and the single choice to make in it; leave every other knob at its default so your numbers match these. The reduction runs on the “Sharp geometry” image; the model runs on the “Clean linear” preset.
1 PCA / SVD (LA 101) -> Image = Sharp geometry, 64x64, seed 1, threshold 95%
2 Regression (Stat 101) -> Preset = Clean linear, seed 3, fit OLS y = b0 + b1 x
Read four numbers across the two tabs: the smallest number of components reaching 95% variance and the reduction ratio it buys in the PCA explorer, then the coefficient of determination and the fitted slope in the regression workshop.
In the PCA / SVD explorer choose the “Sharp geometry” image (defaults: 64×64, no noise, seed 1). This image has 40 non-zero components. Step k up and watch “energy kept” cross 95%; predict the smallest k that reaches it.
Just 3 principal components reach 95% of the variance (EVR = 0.9549) — so the pipeline keeps 3 directions and drops the other 61. This is the demo's own svd() / explained-variance arithmetic, not a typed number.
Keeping k* = 3 of the 64 dimensions is a reduction ratio n / k* = 64 / 3. Predict that ratio — how many times smaller the representation is while it still holds 95% of the variance.
The ratio is 64 / 3 = 21.33× — three principal components stand in for 64 original dimensions. That smaller, denser representation is what feeds the model in stage two. This is the demo's own svd() / explained-variance arithmetic.
Switch to the regression workshop (a Statistics 101 sandbox). Load the “Clean linear” preset with seed 3 and fit ordinary least squares. The coefficient of determination R² = 1 - SSE/SST is the fraction of the response's variance the line explains — the same “variance explained” idea as stage one. Predict whether it is close to 1.
The fit explains R² = 0.9429 of the variance — about 94%, a tight linear relationship. Read against stage one, the parallel is exact: PCA kept 95% of the variance in three directions, and the model captures 94% of the response's. This is the demo's own analyseFit() output, not a typed number.
Same “Clean linear” preset, seed 3. Beyond how well the line fits, its slope b1 is the effect size — how much y moves per unit of x. Predict roughly the slope of the fitted line, then read b1.
The ordinary-least-squares slope is b1 = 1.2508 — each unit of x lifts y by about 1.25. Slope and R² are different facts: the slope is the effect, the R² is how tightly the points hug it. This is the demo's own analyseFit() arithmetic.
Open both sandboxes and run the pipeline yourself: raise the variance threshold and watch how many principal components you must keep grow, then fit the model and watch its R-squared measure the same variance-explained idea on the response side.
A pipeline is a set of stages and the numbers they produce. Write yours down — fill in each blank from the sandboxes, then add one sentence of reasoning. This is the deliverable; there is no number to read off the screen here.
Stage 1 PCA / SVD ...... Sharp geometry, 95% variance k* components .. ______
reduction n/k* . ______ x
Stage 2 Regression ..... Clean linear, seed 3, OLS R-squared ...... ______
slope b1 ....... ______
Pipeline note ........... how does the model's R-squared echo the PCA variance? _____
Design note ............. which knob would you change first, and why? __________
Then change one thing — a higher variance threshold, a noisier regression preset — and note which of the four numbers moved and by how much. That coupling, and the shared “variance explained” language of both stages, is the whole lesson of the project.