Reading
Stories Mode

Types of Data

~12 min read Lesson 2 of 5 in Module 1

Why Data Types Matter

Not all data is created equal. The type of data you collect determines which statistical methods you can use, which visualizations make sense, and what conclusions you can draw. Choosing the wrong method for your data type is one of the most common mistakes in statistics — and it can lead to completely invalid results.

Data types form a hierarchy, from the simplest categories to the richest measurements. Understanding this hierarchy is the essential first step in any statistical analysis. Before you calculate a single number or draw a single chart, you must ask: What kind of data am I working with?

The Golden Rule

The type of data dictates the analysis. You cannot compute a meaningful average of blood types. You cannot rank temperatures in Celsius as ratios. Every statistical tool has assumptions about data type — violating them produces nonsense.

Data Type → Valid Operations → Correct Analysis

Qualitative vs Quantitative

The broadest division in data is between qualitative (categorical) and quantitative (numerical) data. Qualitative data describes qualities or categories — things like colors, names, or labels. Quantitative data represents quantities — things you can count or measure with numbers.

Qualitative data answers what kind? or which category? Quantitative data answers how much? or how many? This distinction is fundamental because the entire toolkit of statistical analysis branches based on it. You summarize categorical data with frequencies and proportions; you summarize numerical data with means and standard deviations.

Nominal Data

Nominal data is the simplest type. It consists of categories with no inherent order. The word “nominal” comes from the Latin nomen, meaning “name” — and that is exactly what nominal data does: it names or labels things.

Examples include blood types (A, B, AB, O), eye colors (brown, blue, green), gender, nationality, and marital status. You can count how many items fall into each category and compute proportions, but you cannot rank these categories or compute a meaningful average. Saying “the average blood type is B+” makes no sense.

Nominal Proportion
\hat{p}_k = \frac{n_k}{N}
The proportion of observations in category k — the only meaningful numerical summary for nominal data

Ordinal Data

Ordinal data has categories with a natural, meaningful order, but the distances between categories are not necessarily equal. Think of customer satisfaction ratings: “very dissatisfied,” “dissatisfied,” “neutral,” “satisfied,” “very satisfied.” There is a clear ranking, but the gap between “neutral” and “satisfied” may not equal the gap between “satisfied” and “very satisfied.”

Other examples include education levels (high school, bachelor's, master's, PhD), military ranks, pain scales (1–10), and socioeconomic status (low, middle, high). You can say one category is “higher” or “better” than another, but you cannot say by how much. Medians and percentiles are appropriate; means are debatable.

Interval Data

Interval data has equal spacing between values, but no true zero point. The classic example is temperature in Celsius or Fahrenheit. The difference between 20°C and 30°C is the same as between 30°C and 40°C — both are 10-degree gaps. But 0°C does not mean “no temperature.” It is an arbitrary reference point (the freezing point of water).

Because there is no true zero, ratios are meaningless for interval data. You cannot say 40°C is “twice as hot” as 20°C. Other examples include calendar dates (year 0 is arbitrary), IQ scores, and standardized test scores. You can compute means and standard deviations, but multiplication and division of values lack interpretation.

Ratio Data

Ratio data is the richest measurement scale. It has equal intervals and a true zero point that represents the complete absence of the quantity. Examples include weight (0 kg means no weight), height, distance, income, age, and time duration.

With ratio data, all mathematical operations are valid. You can say that someone earning $80,000 earns twice as much as someone earning $40,000 — this ratio is meaningful because $0 truly means “no income.” Temperature in Kelvin (where 0 K is absolute zero) is ratio data, even though Celsius and Fahrenheit are not.

Ratio Property
\frac{x_a}{x_b} \text{ is meaningful} \iff \text{true zero exists}
Ratios are meaningful only when a true zero exists — this is what distinguishes ratio from interval data
Measurement Scales Summary
Scale Order Equal Spacing True Zero Example
Nominal No No No Blood type, eye color
Ordinal Yes No No Satisfaction rating
Interval Yes Yes No Temperature (°C)
Ratio Yes Yes Yes Weight, income

Discrete vs Continuous

Within quantitative data, there is another important distinction: discrete vs continuous. Discrete data can only take on specific, countable values — the number of students in a class (you can have 25 or 26, but not 25.7), the number of defective items in a batch, or the count of website visits.

Continuous data can take any value within a range, including fractions and decimals. Height, weight, time, and temperature are all continuous — someone can be 170.3 cm tall, or a reaction can take 3.847 seconds. This distinction matters because discrete and continuous data require different probability distributions and different visualization techniques.

Structured vs Unstructured Data

In the modern data landscape, we also distinguish between structured and unstructured data. Structured data fits neatly into tables with rows and columns — spreadsheets, databases, and CSV files. Each column has a defined data type, and every row follows the same format.

Unstructured data does not have a predefined format. It includes text documents, images, audio recordings, video, social media posts, and sensor streams. An estimated 80–90% of the world's data is unstructured. Processing it requires specialized techniques like natural language processing and computer vision before traditional statistical methods can be applied.

Data Collection Methods

How you collect data is just as important as what type it is. The three primary methods are surveys, experiments, and observational studies.

Surveys collect self-reported data through questionnaires or interviews. They are efficient for gathering large amounts of data but are vulnerable to response bias — people may not answer truthfully or may interpret questions differently.

Experiments are the gold standard for establishing causation. The researcher actively manipulates one variable (the treatment) and measures its effect on another, while controlling for everything else. Randomized controlled trials in medicine are the classic example.

Observational studies collect data without intervening. The researcher watches and records what happens naturally. They can reveal associations and correlations but cannot prove causation — this is one of the most important principles in statistics.

Correlation vs Causation
r_{xy} = \frac{\sum_{i=1}^{n}(x_i - \bar{x})(y_i - \bar{y})}{\sqrt{\sum_{i=1}^{n}(x_i - \bar{x})^2 \cdot \sum_{i=1}^{n}(y_i - \bar{y})^2}}
Correlation measures linear association between two variables — but correlation does not imply causation

Choosing the Right Method

The type of data you have directly determines which statistical methods are appropriate. For nominal data, use frequency tables, bar charts, and chi-squared tests. For ordinal data, use medians, percentiles, and non-parametric tests like the Mann-Whitney U. For interval and ratio data, you have the full arsenal: means, standard deviations, t-tests, ANOVA, regression, and correlation.

Mismatching data types and methods leads to meaningless results. Computing the mean of zip codes, running a t-test on satisfaction ratings, or treating discrete counts as continuous measurements — these are all errors that stem from not understanding your data type. Always classify your data first, then choose your tools.

In the next lesson, we will explore how to visualize data effectively — choosing the right charts and graphs for different data types, and learning to spot patterns, outliers, and distributions at a glance.

Key Takeaways
  • Data falls into two broad categories: qualitative (categorical) and quantitative (numerical) — each requires different analytical approaches.
  • The four measurement scales — nominal, ordinal, interval, and ratio — form a hierarchy of increasing mathematical richness.
  • Discrete data takes countable values; continuous data can take any value in a range — this affects which distributions and tests to use.
  • Data collection method (survey, experiment, observational study) determines what conclusions you can draw — only experiments can establish causation.
  • Always identify your data type before choosing a statistical method — mismatches produce invalid results.
Previous What is Statistics? Module Overview Next Lesson Population vs. Sample