Skip to content

Paired Data

Analyze paired data using the paired t-test and matched pairs designs.

Written and reviewed by the Study Mondo Education TeamLast updated
🎯⭐ INTERACTIVE LESSON

Try the Interactive Version!

Learn step-by-step with practice exercises built right in.

Start Interactive Lesson →

🔗 Tests with Paired Data

When Data Are Paired

Data are paired when observations are linked:

  • Matched subjects: Same person/object measured twice
  • Pre-test/post-test: Before and after intervention
  • Twins or siblings: Natural matching
  • Repeated measures: Same individual under different conditions

Key insight: Pairing reduces variability, improving power of test.

Paired t-Test

When to use: Testing whether mean difference equals zero

Test statistic: t=dˉ−0sd/nt = \frac{\bar{d} - 0}{s_d/\sqrt{n}}

Where:

  • dˉ\bar{d} = mean of differences (for each pair: di=x1i−x2id_i = x_{1i} - x_{2i})
  • sds_d = standard deviation of differences
  • nn = number of pairs
  • Degrees of freedom: df=n−1df = n - 1

Worked Example: A coach tests 12 runners' 100m sprint times before and after training.

RunnerBeforeAfterDifference (dd)
112.412.1-0.3
211.811.5-0.3
............

dˉ=−0.25\bar{d} = -0.25 sec, sd=0.18s_d = 0.18 sec

t=−0.250.18/12=−0.250.052=−4.81t = \frac{-0.25}{0.18/\sqrt{12}} = \frac{-0.25}{0.052} = -4.81

df=11df = 11; t∗=2.201t^* = 2.201 (two-tailed, α=0.05\alpha = 0.05)

Since ∣−4.81∣>2.201|-4.81| > 2.201, reject H0H_0: Training significantly improves sprint time.

Conditions for Paired t-Test

  1. Random sample of pairs
  2. Independence between pairs
  3. Differences approximately normal (or n≥30n \geq 30)

Mean Difference vs. Difference of Means

Do NOT confuse:

  • Mean of differences: dˉ=∑di/n\bar{d} = \sum d_i / n (what we use for paired t-test)
  • Difference of means: xˉ1−xˉ2\bar{x}_1 - \bar{x}_2 (used for two-sample t-test on unpaired data)

For paired data, we always work with the mean of differences.

Common Mistakes

❌ Using two-sample t-test on paired data (loses power) ❌ Computing xˉ1−xˉ2\bar{x}_1 - \bar{x}_2 instead of dˉ\bar{d} ❌ Not pairing the data when possible ❌ Forgetting that nn = number of pairs, not total observations

AP Exam Tip

State "paired t-test" not just "t-test." Define dd clearly: "Let dd = time after minus time before." Show one or two differences calculated.

📚 Practice Problems

1Problem 1easy

❓ Question:

A researcher wants to test if a new study technique improves test scores. She records the scores of 10 students before and after using the technique. Why should she use a paired t-test rather than a two-sample t-test?

💡 Show Solution

She should use a paired t-test because the same students are measured twice (before and after), creating natural pairs. This violates the independence assumption required for two-sample t-tests. The paired design is more powerful because it controls for individual student differences in baseline ability.

Key considerations: • Each student serves as their own control • Focus is on the difference within each pair • Reduces variability by eliminating between-student differences

2Problem 2medium

❓ Question:

Ten married couples were asked to rate their happiness on a scale from 1 to 10. The differences (husband - wife) in ratings were: 2, -1, 0, 3, -2, 1, 0, 2, -1, 1. Construct a 95% confidence interval for the mean difference in happiness ratings.

💡 Show Solution

Step 1: Calculate statistics from differences d̄ = (2 + (-1) + 0 + 3 + (-2) + 1 + 0 + 2 + (-1) + 1) / 10 = 0.5

Step 2: Calculate standard deviation sd = √[Σ(di - d̄)² / (n-1)] = √[14.5 / 9] ≈ 1.27

Step 3: Find t* for df = 9, 95% confidence t* = 2.262

Step 4: Calculate confidence interval CI = d̄ ± t*(sd/√n) CI = 0.5 ± 2.262(1.27/√10) CI = 0.5 ± 0.91 CI = (-0.41, 1.41)

Conclusion: We are 95% confident that the true mean difference in happiness ratings (husband - wife) is between -0.41 and 1.41 points.

3Problem 3medium

❓ Question:

A coach wants to know if a new training program improves 100m sprint times. He records the times of 8 runners before and after the program. The mean difference (before - after) is 0.3 seconds with a standard deviation of 0.4 seconds. Test at α = 0.05 if the program improves times.

💡 Show Solution

H₀: μd = 0 (no improvement) Hₐ: μd > 0 (improvement, before > after)

Test statistic: t = (d̄ - 0) / (sd/√n) t = (0.3 - 0) / (0.4/√8) t = 0.3 / 0.141 t ≈ 2.12

df = n - 1 = 7

P-value (one-tailed): P(t > 2.12) ≈ 0.036

Decision: Since p-value (0.036) < α (0.05), reject H₀

Conclusion: There is sufficient evidence at the 5% significance level to conclude that the training program improves 100m sprint times.

4Problem 4hard

❓ Question:

A pharmaceutical company tests a new medication on 15 patients with high blood pressure. Each patient's blood pressure is measured before treatment and after 3 months. The differences (before - after) have a mean of 8 mmHg and standard deviation of 6 mmHg. Can we conclude at α = 0.01 that the medication lowers blood pressure?

💡 Show Solution

H₀: μd = 0 (no change) Hₐ: μd > 0 (blood pressure decreases)

Test statistic: t = (d̄ - 0) / (sd/√n) t = (8 - 0) / (6/√15) t = 8 / 1.549 t ≈ 5.16

df = 14

P-value (one-tailed): P(t > 5.16) < 0.0001

Decision: Since p-value < 0.01, reject H₀

Conclusion: There is very strong evidence (p < 0.01) that the medication lowers blood pressure. The large t-statistic (5.16) indicates the effect is both statistically significant and likely clinically meaningful.

5Problem 5hard

❓ Question:

A nutritionist studies whether eating breakfast affects students' performance on a math test. She has 20 students take a test after skipping breakfast and another test after eating breakfast (order randomized). Why is this a paired design? What are the advantages and potential concerns?

💡 Show Solution

Why it's paired: Each student takes both tests (no breakfast and with breakfast), creating natural pairs. We analyze the difference in scores for each student.

Advantages: • Controls for individual differences in math ability • More powerful than independent samples design • Requires fewer subjects (20 vs 40 for independent groups) • Each student serves as their own control

Potential concerns:

  1. Practice effect: Students might do better on the second test regardless of breakfast Solution: Randomize which condition comes first

  2. Carryover effect: Effects from first test might influence second test Solution: Sufficient time between tests

  3. Different test difficulty: If tests aren't equivalent, this confounds results Solution: Use equivalent forms or counterbalance test versions

  4. Learning between tests: Students might study between tests Solution: Control time between tests, avoid giving feedback

Explain using:

⚠️ Common Mistakes: Paired Data

Avoid these 3 frequent errors

📌 Related Topics in Unit 7: Inference for Quantitative Data — Means

❓ Frequently Asked Questions

What is Paired Data?▾
Analyze paired data using the paired t-test and matched pairs designs.
How can I study Paired Data effectively?▾
Start by reading the study notes and working through the examples on this page. Then use the flashcards to test your recall. Practice with the 5 problems provided, checking solutions as you go. Regular review and active practice are key to retention.
Is this Paired Data study guide free?▾
Yes — all study notes, flashcards, and practice problems for Paired Data on Study Mondo are free to access. No account is needed.
What course covers Paired Data?▾
Paired Data is part of the AP Statistics course on Study Mondo, specifically in the Unit 7: Inference for Quantitative Data — Means section. You can explore the full course for more related topics and practice resources.
Are there practice problems for Paired Data?▾
Yes, this page includes 5 practice problems with detailed solutions. Each problem includes a step-by-step explanation to help you understand the approach.