Inference for Two Sample Proportions (CI and Test) - Complete Interactive Lesson
Part 1: Two-Sample Z-Test for Proportions
⚖️ Comparing Two Populations
Part 1 of 7 — Two-Sample Z-Test for Proportions
When to Compare Two Proportions
Use when you have two independent groups and want to test whether their population proportions differ.
Example: Is the proportion of smartphone users higher among teens than adults?
Hypotheses
Test Statistic
where is the pooled proportion.
Conditions
- Random samples from both populations
- Independent groups (and of population for each)
- Large counts:
🔑 Under , we assume , so we use the pooled proportion for the standard error.
Two-Proportion Test 🎯
Two-Proportion Calculations 🧮
Treatment group: 45 successes out of 150. Control group: 30 successes out of 150.
1) ? (Express as a decimal)
2) ?
3) Pooled ?
Part 2: Two-Sample T-Test for Means
📊 Two-Sample T-Test for Means
Part 2 of 7 — Comparing Two Independent Groups
Topics in This Part
| Section |
|---|
| 🎯 Hypotheses for |
| 📐 The Two-Sample -Statistic |
| ✅ Conditions |
| 📝 Worked Example |
🔑 Key Concept: The two-sample -test compares means from two independent groups. The key question: "Is the difference in sample means large enough to conclude the population means differ?"
The Setup
Two independent random samples:
- Group 1: observations, sample mean , sample SD
- Group 2: observations, sample mean , sample SD
Hypotheses
Test Statistic
Degrees of freedom: Use the calculator/technology value (Welch's approximation) or the conservative value .
⚠️ AP Tip: The AP formula sheet provides this formula. On the AP exam, use the conservative df or the calculator df — both are acceptable.
Conditions
| Condition | Check |
|---|---|
| Random | Both samples are random (or randomly assigned in an experiment) |
| Independent | The two groups are independent of each other |
| Normal/Large Sample | Both populations are Normal OR both and |
| 10% Rule | of population 1 AND of population 2 |
Worked Example
Study: Does a new teaching method improve test scores?
- Method A (traditional): , ,
- Method B (new): , ,
Step 1 — Hypotheses: (no difference in mean scores) (new method produces higher mean scores)
Step 2 — Conditions:
- Random: Students randomly assigned to methods ✓
- Independent: Two separate groups ✓
- Normal: and ✓
- 10%: Each group is less than 10% of all students ✓
Step 3 — Test statistic:
df (conservative) or use calculator df
Step 4 — Conclusion: (one-sided). Since , we reject . There is convincing evidence that the new teaching method produces higher mean test scores than the traditional method.
Two-Sample -Test Concepts 🎯
Computing the -Statistic 🧮
Group 1: , , Group 2: , ,
1)
2) (compute to one decimal)
3) Conservative df
Decisions and Interpretations 🔍
Exit Quiz — Two-Sample -Test ✅
Part 3: Paired T-Test
📊 Paired T-Test
Part 3 of 7 — When Data Come in Pairs
Topics in This Part
| Section |
|---|
| 🔗 Paired vs. Two-Sample Designs |
| 📐 The Paired -Test Procedure |
| ✅ Conditions for Paired Data |
| 📝 Worked Example |
🔑 Key Concept: A paired -test is a one-sample -test performed on the differences within each pair.
When to Use a Paired Test
Use a paired -test when:
- The same subjects are measured twice (before/after)
- Subjects are matched in pairs (e.g., twins, partners)
- Each observation in one group has a natural partner in the other
Use a two-sample -test when:
- The two groups are independent (different subjects, no pairing)
The Procedure
Step 1: Compute the differences: for each pair
Step 2: Treat the differences as a single sample and perform a one-sample -test:
where:
- = mean of the differences
- = standard deviation of the differences
- = number of pairs
- df
Hypotheses
⚠️ AP Tip: Define clearly: " = After Before" or " = Treatment Control." The direction matters for one-sided tests.
Conditions
| Condition | Check |
|---|---|
| Random | Pairs are randomly selected or treatments randomly assigned |
| Independent | Individual pairs are independent of each other (10% condition on pairs) |
| Normal | The differences () are approximately Normal (check histogram/QQ plot of ) |
Worked Example
Study: Does a tutoring program improve SAT scores? 12 students take the SAT, receive tutoring, then retake it.
| Student | Before | After | = After Before |
|---|---|---|---|
| 1 | 520 | 560 | |
| 2 | 480 | 510 | |
| 3 | 550 | 540 | |
| ... | ... | ... | ... |
| Summary | , |
Step 1 — Hypotheses: (tutoring has no effect on mean SAT scores) (tutoring increases mean SAT scores) where = After Before
Step 2 — Conditions:
- Random: Students randomly selected ✓
- Independent: 12 students < 10% of all SAT takers ✓
- Normal: Histogram of differences shows no strong skew ✓
Step 3 — Test statistic: df
Step 4 — Conclusion: (one-sided). Since , we reject . There is convincing evidence that the tutoring program increases mean SAT scores.
Paired vs. Two-Sample: Why It Matters
| Feature | Two-Sample | Paired |
|---|---|---|
| Data structure | Two independent groups | Pairs of related observations |
| Parameter | ||
| Advantage | Simpler design | Controls for variability between subjects |
| df | Complex (Welch) |
🔑 Key Insight: Pairing reduces variability by eliminating subject-to-subject differences, making it easier to detect a treatment effect.
Paired -Test Concepts 🎯
Computing the Paired 🧮
pairs, ,
1) SE of
2)
3) df
Paired or Two-Sample? 🔍
Exit Quiz — Paired -Test ✅
Part 4: Confidence Intervals for Differences
📊 Confidence Intervals for Differences
Part 4 of 7 — Estimating How Much Two Groups Differ
Topics in This Part
| Section |
|---|
| 📐 Two-Sample CI for |
| 🔗 Paired CI for |
| 📝 Interpretation Templates |
| 🧮 Connection to Hypothesis Tests |
🔑 Key Concept: A CI for the difference gives a range of plausible values for how much two population means (or the mean difference) differ.
Two-Sample CI for
- Same conditions as the two-sample -test (Random, Independent, Normal/Large, 10%)
- df: use calculator (Welch) or conservative
Paired CI for
- Same conditions as the paired -test (Random, Independent pairs, Normal differences)
- df (where = number of pairs)
Interpretation Templates
Two-Sample: "We are [C]% confident that the true difference in mean [context] between [group 1] and [group 2] is between [lower] and [upper] [units]."
Paired: "We are [C]% confident that the true mean difference in [context] is between [lower] and [upper] [units]."
CI ↔ Test Connection
| CI contains 0? | Test conclusion at |
|---|---|
| Yes | Fail to reject (or ) |
| No | Reject |
Worked Example — Two-Sample CI
Group A (old drug): , , Group B (new drug): , ,
95% CI for :
- SE
- Conservative df ,
Interpretation: "We are 95% confident that the true difference in mean recovery scores (new old) is between 1.1 and 10.9 points. Since 0 is not in the interval, there is evidence the new drug produces higher mean scores."
Worked Example — Paired CI
15 patients measured before and after treatment. (After Before),
95% CI for : df ,
Interpretation: "We are 95% confident that the true mean change in [outcome] after treatment is between 4.9 and 11.7 units."
⚠️ AP Tip: Always define (e.g., After Before) and include context and units.
CI for Differences Concepts 🎯
Building CIs 🧮
Two-Sample: , SE ,
1) Margin of error
2) Lower bound
Paired: , , ,
3) SE of
Interpretation Decisions 🔍
Exit Quiz — CIs for Differences ✅
Part 5: Power and Sample Size
📊 Power and Sample Size
Part 5 of 7 — Detecting Real Differences
Topics in This Part
| Section |
|---|
| ⚡ What Is Power? |
| 🎯 Type I and Type II Errors |
| 📐 Factors Affecting Power |
| 🧮 Sample Size Considerations |
🔑 Key Concept: Power is the probability of correctly rejecting when is actually false. Higher power = better ability to detect a real effect.
Error Types
| True | False | |
|---|---|---|
| Reject | Type I Error () | Correct! (Power) |
| Fail to Reject | Correct! | Type II Error () |
Type I Error ()
- Rejecting when it is true (false positive)
- Probability (the significance level)
- Example: Concluding a drug works when it actually does not
Type II Error ()
- Failing to reject when it is false (false negative)
- Probability =
- Example: Concluding a drug does not work when it actually does
Factors That Increase Power
| Factor | Direction | Effect on Power |
|---|---|---|
| Sample size () | ↑ | Power ↑ |
| Significance level () | ↑ | Power ↑ |
| True effect size ($ | \mu_1 - \mu_2 | $) |
| Population variability () | ↓ | Power ↑ |
⚠️ AP Tip: You will NOT be asked to calculate power on the AP exam, but you MUST understand conceptually how each factor affects power.
Intuition for Each Factor
Larger : More data → smaller SE → easier to detect a difference
Larger : Easier rejection threshold → more likely to reject (but more risk of Type I error)
Larger effect: A bigger real difference is easier to detect than a tiny one
Smaller : Less noise → the signal (difference) stands out more clearly
The Power- Tradeoff
Decreasing (e.g., from 0.05 to 0.01) reduces Type I error but increases Type II error (reduces power). The only way to reduce BOTH errors simultaneously is to increase sample size.
Sample Size Planning
Before collecting data, researchers choose to achieve desired power (typically 80% or higher):
- Specify the smallest meaningful effect size
- Estimate population variability ()
- Choose (usually 0.05)
- Use a power table or software to find the required
🔑 Key Insight: Larger samples are always better for power, but they cost more. Sample size planning balances statistical needs with practical constraints.
Power Concepts 🎯
Error and Power Calculations 🧮
1) If , what is the power? (give as decimal)
2) If , what is the probability of a Type I error?
3) A test has power . What is ?
Power Factors 🔍
Exit Quiz — Power & Sample Size ✅
Part 6: Problem-Solving Workshop
📊 Problem-Solving Workshop
Part 6 of 7 — Complete Worked Examples
Worked Example 1: Two-Sample T-Test
Problem
A fitness company wants to compare two training programs. 45 volunteers are randomly assigned: 22 to Program A, 23 to Program B. After 8 weeks, weight loss (in pounds) is recorded:
| Program A | Program B | |
|---|---|---|
| 22 | 23 | |
| 12.4 | 9.7 | |
| 3.8 | 4.1 |
Test whether Program A produces greater average weight loss at .
Step 1: State Hypotheses
Where = true mean weight loss for Program A, = true mean weight loss for Program B.
Step 2: Check Conditions
✅ Random: Volunteers randomly assigned to groups (experiment) ✅ Normal: and (or no strong skewness mentioned) ✅ Independent: Groups are independent; each person in only one program. Both of all potential participants.
Step 3: Calculate
Using technology with : -value
Step 4: Conclude
Since , we reject .
AP-Style Conclusion: There is convincing evidence that the true mean weight loss for Program A is greater than the true mean weight loss for Program B.
Worked Example 2: Paired T-Test
Problem
A researcher tests whether a meditation app reduces stress. 30 participants rate their stress (1–100) before and after 4 weeks of daily use:
| Before | After | Differences (Before − After) | |
|---|---|---|---|
| 30 | 30 | 30 | |
| 68.2 | 59.5 | ||
| — | — |
Test whether the app reduces stress at .
Step 1: State Hypotheses
Where = true mean difference in stress scores (Before − After) for all users of this app.
Step 2: Check Conditions
✅ Random: 30 participants randomly selected (or assume representative) ✅ Normal: (CLT applies for differences) ✅ Independent: Differences within each person are independent; of all potential users
⚠️ Key: We check conditions on the DIFFERENCES, not the individual scores.
Step 3: Calculate
, -value
Step 4: Conclude
Since , we reject .
AP-Style Conclusion: There is convincing evidence that the meditation app reduces mean stress scores.
Common AP Mistakes
| Mistake | Why It Costs Points |
|---|---|
| Using two-sample test when data is paired | Wrong procedure → wrong test statistic → wrong conclusion |
| Not defining clearly | "Mean difference" must include direction (A − B) and context |
| Skipping conditions | Automatic deduction on free-response |
| No context in conclusion | Must reference the specific variables and setting |
| Saying "accept " | Always say "fail to reject " |
| Not identifying data as paired | Look for: same subjects, before/after, matched pairs |
Workshop Practice 🎯
Computation Practice 🧮
1) Two groups: , , ; , , . Calculate SE (round to 2 decimal places).
2) Using the values above, calculate the t-statistic (round to 2 decimal places).
3) Paired data: , , . Calculate the t-statistic (round to 2 decimal places).
Procedure Selection 🔍
Exit Quiz — Problem-Solving Workshop ✅
Part 7: Review & Applications
📊 Review & Applications
Part 7 of 7 — Comprehensive Review
Complete Formula Reference
Two-Sample T-Test for Means
| Component | Formula |
|---|---|
| Standard Error | |
| Test Statistic | |
| Confidence Interval | |
| Degrees of Freedom | Use technology (Welch's approximation) |
Paired T-Test
| Component | Formula |
|---|---|
| Mean Difference | |
| Standard Error | |
| Test Statistic | |
| Confidence Interval | |
| Degrees of Freedom |
Decision Guide: Paired vs. Two-Sample
| Question | Paired | Two-Sample |
|---|---|---|
| Same subjects measured twice? | ✅ | ❌ |
| Before/after design? | ✅ | ❌ |
| Matched pairs (twins, siblings)? | ✅ | ❌ |
| Two independent groups? | ❌ | ✅ |
| Random assignment to groups? | ❌ | ✅ |
| Check conditions on... | Differences | Each sample |
| Technology |
Conditions Summary
Two-Sample Tests/CIs
- Random: Both samples from random processes (or random assignment)
- Normal: and , or no strong skewness/outliers
- Independent: Samples independent of each other; each of its population
Paired Tests/CIs
- Random: Random sample of pairs (or randomly determine order)
- Normal: differences, or differences show no strong skewness/outliers
- Independent: Individual pairs are independent; of all pairs in population
Interpretation Templates
Hypothesis Test Conclusion
"Since -value = ___ is [less/greater] than ___, we [reject/fail to reject] . There [is/is not] convincing evidence that [context: what the alternative hypothesis claims]."
Confidence Interval Interpretation
"We are ___% confident that the true difference in means [context: in words] is between ___ and ___."
Confidence Interval and Significance
- If 0 is NOT in the CI → Reject (at the corresponding )
- If 0 IS in the CI → Fail to reject
Key Concepts from Every Part
| Part | Topic | Key Idea |
|---|---|---|
| 1 | Introduction | Two-sample vs. paired designs |
| 2 | Two-Sample T-Test | Tests for differences between independent groups |
| 3 | Paired T-Test | Tests for differences within matched pairs |
| 4 | Confidence Intervals | Estimate the true difference with a range |
| 5 | Power & Sample Size | Power = ; larger → more power |
| 6 | Problem-Solving | Complete 4-step process with real data |
Power Quick Review
Power increases when: , , effect size ,
🔑 Final Tip: The AP exam tests your ability to (1) choose the right test, (2) check conditions, (3) calculate correctly, and (4) interpret in context. Practice the full 4-step process until it becomes automatic.
Comprehensive Review 🎯
Formula Application 🧮
1) Two-sample data: , , . Calculate (round to 1 decimal place).
2) Paired data: , , . Calculate (round to 1 decimal place).
3) Using the values from #2, what is ?
Concept Connections 🔍
Final Exam — Comparing Populations ✅