Types of Data and Sampling
Learn to identify categorical vs. quantitative data, and understand different sampling methods.
Try the Interactive Version!
Learn step-by-step with practice exercises built right in.
Types of Data and Sampling
Types of Data
Categorical (Qualitative) Data
- Data that describes qualities or characteristics
- Cannot be ordered meaningfully (not about size/amount)
- Examples: color, political party, type of fruit
- Divided into:
- Nominal: no natural ordering (red/blue/green)
- Ordinal: natural ordering (small/medium/large, freshman/sophomore/junior/senior)
Quantitative (Numerical) Data
- Data that represents measurements or amounts
- Can be ordered and arithmetic makes sense
- Examples: height, weight, age, test score
- Divided into:
- Discrete: counts, whole numbers (number of siblings, cars sold)
- Continuous: measured, any value in a range (height, time, weight)
Key Definitions
Population: entire group of individuals we want information about
- Example: all AP Statistics students in the United States
Sample: subset of the population we actually collect data from
- Example: 500 randomly selected AP Statistics students
Parameter: numerical summary of a population (unknown, fixed)
- Notation: usually Greek letters (μ, σ, p)
- Example: true mean score of all AP Stat students
Statistic: numerical summary of a sample (known, varies sample to sample)
- Notation: usually letters (\(ar{x}\), s, \(hat{p}\))
- Example: mean score of 500 sampled students
Sampling Methods
Simple Random Sample (SRS)
- Every individual has equal chance of selection
- Best method when population is accessible
- Use random number generator or table
- Eliminates selection bias
Stratified Random Sample
- Divide population into homogeneous groups (strata)
- Randomly sample from each stratum proportionally
- Better representation of subgroups
- Example: stratify by grade level, then randomly select from each
Cluster Sample
- Divide population into clusters (typically geographic)
- Randomly select entire clusters
- Less expensive than SRS; useful when population spread out
- Risk: clusters not representative
Systematic Sample
- Select every \(k\)-th individual from ordered list
- Simple to implement
- Risk: hidden pattern in list
Convenience Sample (avoid for inference)
- Sample whoever is easiest to reach
- Biased; generally not representative
Common Sampling Biases
Sampling bias (selection bias): certain individuals more likely to be selected
- Convenience sample at mall (misses online shoppers)
Response bias: individuals respond untruthfully or refuse
- Loaded question: "Don't you agree this policy is wasteful?"
- Shy respondents not answering honestly
Non-response bias: some selected individuals don't respond
- Mail survey with 40% return rate
Undercoverage: some part of population not accessible
- Phone survey (misses homeless)
Worked Example
Scenario: A school wants to estimate mean SAT score for all 1,200 seniors.
Method 1 (SRS): Generate 100 random student IDs from 1–1200, compute mean score for those students.
Method 2 (Stratified): Divide into 3 strata by gender (400 male, 500 female, 300 nonbinary). From each stratum, randomly select 33–34 students. Compute mean.
Method 3 (Systematic): Generate random starting point (say, 5), then select students 5, 17, 29, 41, ... until 100 selected.
Decision Rule: When to Use Each
- SRS: population list available, want unbiased sample, resources sufficient
- Stratified: important subgroups exist, want equal representation of subgroups
- Cluster: population geographically spread, population list unavailable, budget limited
- Avoid convenience/systematic: if inference accuracy is important
AP Exam Tip
On FRQ prompt about study design, identify:
- Population (who are we studying?)
- Sample method (how were subjects selected?)
- Bias present? (selection, response, non-response, undercoverage?)
- Why this method? (explain trade-offs)
Common error: assuming \(ar{x}\) = μ just because you have a large sample. Sampling bias can produce bad estimates even with large n.
📚 Practice Problems
1Problem 1easy
❓ Question:
A survey asks students: 'Do you prefer morning or afternoon classes?' and 'How many hours per week do you study?' Classify each variable.
💡 Show Solution
The first variable (morning or afternoon) is categorical because it describes a category preference with no numerical value. The second variable (study hours) is quantitative (specifically continuous) because it represents a measurable quantity that can take any value within a range. In data collection, categorical variables describe qualities while quantitative variables measure quantities.
2Problem 2medium
❓ Question:
A researcher wants to estimate the average GPA of all 10,000 students at a university. She randomly selects 250 students and calculates their mean GPA as 3.42. Identify the population, sample, parameter, and statistic.
💡 Show Solution
Population: All 10,000 students at the university (the entire group of interest)
Sample: The 250 randomly selected students (the subset actually studied)
Parameter: The average (mean) GPA of all 10,000 students — this is unknown and what the researcher wants to estimate. It's a fixed value describing the population.
Statistic: The mean GPA of 3.42 from the sample of 250 students — this is known from the data and used to estimate the parameter.
Key distinction: A parameter describes a population (usually unknown); a statistic describes a sample (known from data) and is used to estimate the parameter.
3Problem 3hard
❓ Question:
Explain why a census might be impractical for estimating the average lifespan of light bulbs manufactured by a company, and explain what sampling method you would use instead.
💡 Show Solution
Why a census is impractical:
A census requires testing all light bulbs produced by the company. This would mean destroying them all to measure their lifespan — the manufacturer would have no product left to sell! This is both destructive and economically infeasible.
Better approach: Sampling
Use random sampling (specifically, a simple random sample or SRS) by:
- Randomly selecting a representative subset of light bulbs from the entire production
- Testing these to destruction and recording lifespans
- Computing the sample mean lifespan
- Using this to estimate the population parameter
This preserves most inventory, is cost-effective, and gives a reliable estimate when the sample size is adequate. The randomness ensures the sample is representative of all bulbs produced.
⚠️ Common Mistakes: Types of Data and Sampling
Avoid these 3 frequent errors
Practice with Flashcards
Rate this topic's cards with spaced repetition. Cards join your deck when you finish a topic's lesson and take its exit quiz.
Browse All Topics
Explore more AP Statistics topics