Skip to content

Describing Distributions

Describe the shape, center, spread, and outliers of a distribution using SOCS.

Written and reviewed by the Study Mondo Education TeamLast updated
🎯⭐ INTERACTIVE LESSON

Try the Interactive Version!

Learn step-by-step with practice exercises built right in.

Start Interactive Lesson →

Describing Distributions

SOCS Framework

Always describe a distribution using S–O–C–S:

  1. Shape
  2. Outliers
  3. Center
  4. Spread

Shape

Symmetry:

  • Symmetric: left and right halves mirror each other (mean ≈ median)
  • Skewed left (negatively skewed): tail extends left, peak right (mean < median)
  • Skewed right (positively skewed): tail extends right, peak left (mean > median)

Modality:

  • Unimodal: one peak
  • Bimodal: two peaks (often two subpopulations)
  • Multimodal: more than two peaks
  • Uniform: roughly equal height across range

Peakedness:

  • Roughly normal (bell curve)
  • Flatter than normal (platykurtic)
  • Sharper than normal (leptokurtic)

Outliers

Definition: observations unusually far from the rest

  • Identify using boxplot (beyond whiskers using 1.5·IQR rule)
  • Or context: "100 hours of TV watching" when most watch <20

Impact:

  • Pull mean toward outlier (mean not resistant)
  • Median unaffected (median is resistant)
  • Increase standard deviation and range

Investigation: is it a genuine measurement, data entry error, or unusual case?

Center

Mean (\(\bar{x}\)): arithmetic average

  • \(\bar{x} = \frac{\sum x_i}{n}\)
  • Pulled by outliers (not resistant)

Median: middle value when ordered

  • 50th percentile
  • Resistant to outliers
  • Preferred for skewed distributions

Mode: most frequent value

  • Used for categorical or discrete data

Rule of thumb:

  • Symmetric distribution: mean ≈ median
  • Skewed distribution: prefer median

Spread

Range: max − min

  • Simplest measure
  • Affected by outliers
  • Not resistant

Interquartile Range (IQR): \(Q3 - Q1\)

  • Middle 50% of data
  • Resistant to outliers
  • Preferred for skewed distributions

Variance (\(s^2\)): average squared deviation from mean

  • \(s^2 = \frac{\sum(x_i - \bar{x})^2}{n-1}\) (sample variance, divide by n−1)

Standard Deviation (\(s\)): square root of variance

  • \(s = \sqrt{s^2}\)
  • Same units as data
  • Measures typical distance from mean
  • Estimated: in roughly normal data, about 68% within 1s of mean

Worked Example

Data: Heights (inches) of 10 students: 62, 64, 65, 66, 67, 68, 69, 71, 72, 75

SOCS Description:

  1. Shape: roughly unimodal and symmetric (slight right skew due to 75)
  2. Outliers: boxplot Q1 ≈ 65.5, Q3 ≈ 70.5, IQR = 5; fences at 65.5 − 7.5 = 58 and 70.5 + 7.5 = 78. No outliers.
  3. Center: mean = \(\frac{62+64+...+75}{10} = 67.9\) inches; median = \(\frac{67+68}{2} = 67.5\) inches (very close, confirming near symmetry)
  4. Spread: range = 75 − 62 = 13 inches; IQR = 5 inches; \(s \approx 3.7\) inches

Comparison Language

When comparing two distributions:

Shape: "Distribution A is symmetric while Distribution B is right-skewed."

Center: "The median for Group X is approximately _____ inches, compared to _____ inches for Group Y, so Group X tends to be taller."

Spread: "Group X has an IQR of _____, while Group Y has IQR of _____, so Group Y is more variable."

Outliers: "Distribution A has one outlier at _____, while Distribution B has no outliers."

Common Mistakes

  1. Saying "mean = 50" when you haven't calculated it
  2. Forgetting to identify shape when asked to describe
  3. Confusing resistant vs. non-resistant (median is resistant; mean is not)
  4. Using mean and median interchangeably in skewed data

AP Exam Tip

On free response, examiners want to see you use SOCS explicitly. Write:

  • "Shape: ..."
  • "Outliers: ..."
  • "Center: ..."
  • "Spread: ..."

Use appropriate statistics for the shape (median/IQR for skewed; mean/std dev for symmetric).

📚 Practice Problems

1Problem 1easy

❓ Question:

A distribution shows most values clustered near 50, with a long tail extending to the right toward 100. Describe the shape and identify where the mean is relative to the median.

💡 Show Solution

This distribution is skewed to the right (positively skewed). When data is skewed right, the mean is pulled toward the tail (toward the higher values), so the mean > median. The long tail of extreme high values increases the average more than it affects the median. This is common in real-world data like incomes or test scores with a ceiling effect.

2Problem 2medium

❓ Question:

Two datasets have the same median (both 70) but different shapes: Dataset A is symmetric, while Dataset B is skewed left. Without seeing the distributions, explain what this tells you about their means.

💡 Show Solution

Dataset A (symmetric): Mean ≈ Median ≈ 70, because symmetry means the values balance equally on both sides.

Dataset B (skewed left): Mean < Median. The median is 70, but the mean is pulled toward the left tail by the extreme low values. The mean will be less than 70.

Why? In a left-skewed distribution, the tail extends toward lower values. These extreme low outliers pull the mean down more than they affect the median (which is just the middle position). This is common in age-at-death data or grade distributions where there's a floor but not a ceiling.

3Problem 3hard

❓ Question:

A histogram shows test scores for a large class with two distinct peaks: one at 70 and another at 85. Interpret this distribution and suggest what might explain it. How would you describe center and spread?

💡 Show Solution

Distribution characteristics:

This is a bimodal distribution with two peaks (modes) at 70 and 85. The presence of two modes suggests two distinct groups within the class, not a single homogeneous population.

Possible explanations:

  • Different preparation levels (some students studied more thoroughly than others)
  • One mode represents students who barely passed; the other represents strong performers
  • The class might contain different ability levels mixed together

Center & Spread:

  • A single measure of center (like the mean) would be misleading—it would fall around 77-78, not representing either peak well
  • Better approach: Report both modes separately, or note that the distribution is bimodal
  • Spread: The distance between the two peaks (15 points) is noteworthy; report the range and/or standard deviation, noting the gap suggests distinct subgroups

Lesson: Always look for multiple peaks; they suggest mixture of populations.

Explain using:

⚠️ Common Mistakes: Describing Distributions

Avoid these 3 frequent errors

📌 Related Topics in Unit 1: Exploring One-Variable Data

❓ Frequently Asked Questions

What is Describing Distributions?▾
Describe the shape, center, spread, and outliers of a distribution using SOCS.
How can I study Describing Distributions effectively?▾
Start by reading the study notes and working through the examples on this page. Then use the flashcards to test your recall. Practice with the 3 problems provided, checking solutions as you go. Regular review and active practice are key to retention.
Is this Describing Distributions study guide free?▾
Yes — all study notes, flashcards, and practice problems for Describing Distributions on Study Mondo are free to access. No account is needed.
What course covers Describing Distributions?▾
Describing Distributions is part of the AP Statistics course on Study Mondo, specifically in the Unit 1: Exploring One-Variable Data section. You can explore the full course for more related topics and practice resources.
Are there practice problems for Describing Distributions?▾
Yes, this page includes 3 practice problems with detailed solutions. Each problem includes a step-by-step explanation to help you understand the approach.