Reading
Stories Mode

What is Statistics?

~10 min read Lesson 1 of 5 in Module 1

The Science of Learning from Data

Statistics is the science of learning from data. It provides us with the tools to collect, organize, analyze, and interpret information — helping us make decisions under uncertainty. From clinical trials that determine whether a new drug works, to election polls that predict voter behavior, to A/B tests that optimize your favorite apps, statistics is everywhere.

At its core, statistics answers a simple but powerful question: What can we conclude from the data we have, and how confident should we be in that conclusion? This question sits at the heart of science, business, medicine, and public policy.

Why Statistics Matters

We live in a world of incomplete information. We can rarely observe everything, measure everyone, or know the future with certainty. Statistics gives us a rigorous framework to make sense of what we can observe and draw reliable conclusions from it.

Data + Statistical Methods = Informed Decisions Under Uncertainty
Timeline — The Birth of Statistics
1654
Pascal & Fermat develop probability theory through correspondence about gambling problems
1809
Gauss publishes the normal distribution — the bell curve that governs countless natural phenomena
1900
Pearson introduces the chi-squared test, enabling rigorous hypothesis testing
1925
Fisher publishes Statistical Methods for Research Workers — modern statistics is born

Data and Uncertainty

We live in a world of incomplete information. Every measurement has some noise, every sample is only a fraction of the whole, and the future is never certain. Statistics gives us tools to quantify uncertainty and make decisions despite it.

Think about a doctor deciding whether a treatment works. They cannot test every patient on Earth — they test a sample and use statistics to determine whether the observed improvement is real or just due to random chance. This is the fundamental challenge that statistics addresses.

Population vs Sample

One of the most important distinctions in statistics is between a population and a sample. The population is the entire group you want to study — every voter in a country, every smartphone in production, every patient with a disease. The sample is the subset you actually observe.

We almost never have access to the full population. Instead, we collect a sample and use it to make inferences about the population. The art and science of statistics lies in knowing how to draw valid conclusions from limited data.

The Key Distinction

Population: The complete set of all elements you want to study (often too large to measure entirely).

Sample: A subset of the population that you actually collect data from. Good samples are representative of the population.

Descriptive Statistics

Before making any inferences, we first need to summarize our data. Descriptive statistics are the tools that help us do this. They answer the question: What does our data look like?

The most common descriptive statistics include measures of central tendency (where is the center of the data?) and measures of spread (how variable is the data?). The mean, median, and mode tell us about the center; variance and standard deviation tell us about spread.

The Sample Mean
\bar{x} = \frac{1}{n}\sum_{i=1}^{n}x_i
The arithmetic mean — the sum of all values divided by the number of observations
Sample Standard Deviation
s = \sqrt{\frac{1}{n-1}\sum_{i=1}^{n}(x_i - \bar{x})^2}
Measures how spread out the data is around the mean — low values mean data clusters tightly, high values mean wide spread. The divisor is n − 1, not n: one degree of freedom is spent estimating the sample mean, and Module 5 shows why that correction makes s² an unbiased estimate of the population variance σ².

Inferential Statistics

While descriptive statistics summarize what we see, inferential statistics let us draw conclusions about what we cannot see. Using data from a sample, we make claims about the larger population.

The two main tools of inferential statistics are hypothesis testing and confidence intervals. Hypothesis testing asks: Is there enough evidence to support a claim? Confidence intervals give us a range of plausible values for a population parameter. Together, they form the backbone of scientific research.

Probability: The Foundation

Underlying all of statistics is probability — the mathematical language of uncertainty. Probability assigns a number between 0 and 1 to every possible outcome: 0 means impossible, 1 means certain, and values in between reflect degrees of belief or likelihood.

Classical Probability
P(A) = \frac{\text{favorable outcomes}}{\text{total outcomes}}
The probability of an event A is the ratio of favorable outcomes to the total number of equally likely outcomes

Probability is the thread that connects data to decisions. Every statistical test, every confidence interval, every prediction model is built on the foundation of probability theory.

Statistics Everywhere

Statistics is not confined to textbooks. It is a living, breathing part of the modern world. Here are just a few areas where statistical thinking drives real decisions:

Medicine: Clinical trials use statistics to determine whether a drug is effective and safe. Without statistical rigor, we cannot distinguish real treatments from placebos.

Technology: A/B testing — the backbone of product decisions at companies like Google, Netflix, and Amazon — is applied statistics. Every button color, recommendation algorithm, and pricing strategy is tested statistically.

Sports: From Moneyball to modern analytics, teams use statistics to evaluate players, optimize strategies, and gain competitive advantages.

Government: Census data, economic indicators, and public health surveillance all depend on statistical methods to inform policy decisions.

Finance: Risk assessment, portfolio optimization, and insurance pricing are all fundamentally statistical problems.

In the next lesson, we will dive into types of data and how to visualize them — the essential first step in any statistical analysis. You will learn about categorical vs. numerical data, distributions, and how to choose the right chart for your data.

Key Takeaways
  • Statistics is the science of learning from data under uncertainty — it helps us make informed decisions when we cannot know everything.
  • Population (everyone) vs Sample (a subset) is the key distinction — we almost always work with samples and infer about populations.
  • Descriptive statistics summarize data (mean, median, standard deviation); inferential statistics make predictions and test hypotheses.
  • Probability is the mathematical language of uncertainty — it is the foundation on which all of statistics is built.
Previous None Module Overview Next Lesson Types of Data