Skip to content
🎯⭐ INTERACTIVE LESSON

Chi-Square Tests

Learn step-by-step with interactive practice!

Chi-Square Tests - Complete Interactive Lesson

Part 1: Chi-Square Goodness-of-Fit

📊 Chi-Square Goodness-of-Fit Test

Part 1 of 7 — Testing Categorical Distributions


When to Use Chi-Square Goodness-of-Fit

Use when you want to test whether observed frequencies match expected frequencies for a categorical variable.

Example: A die is rolled 60 times. Do the results suggest it is fair?

Outcome123456
Observed81271599
Expected101010101010

The Chi-Square Statistic

χ2=∑(Oi−Ei)2Ei\chi^2 = \sum \frac{(O_i - E_i)^2}{E_i}

where OiO_i = observed count and EiE_i = expected count.

For the die example: χ2=(8−10)210+(12−10)210+⋯=4+4+9+25+1+110=4.4\chi^2 = \frac{(8-10)^2}{10} + \frac{(12-10)^2}{10} + \cdots = \frac{4+4+9+25+1+1}{10} = 4.4


Hypotheses

  • H0H_0: The observed distribution matches the expected distribution
  • HaH_a: The observed distribution does NOT match the expected

Conditions

  1. Random sample or random assignment
  2. Expected counts ≥ 5 for all categories
  3. Independence — observations are independent

🔑 Chi-square tests are always right-tailed — larger χ2\chi^2 values provide more evidence against H0H_0.

Goodness-of-Fit Check 🎯

Chi-Square Calculation 🧮

A bag should contain equal numbers of 4 colors. From 80 candies: Red=24, Blue=18, Green=22, Yellow=16.

1) Expected count for each color = 80/4 = ?

2) Compute χ2=(24−20)220+(18−20)220+(22−20)220+(16−20)220\chi^2 = \frac{(24-20)^2}{20} + \frac{(18-20)^2}{20} + \frac{(22-20)^2}{20} + \frac{(16-20)^2}{20}

3) Degrees of freedom = k−1k - 1 = ?

Part 2: Chi-Square Test for Independence

📊 Chi-Square Test for Independence

Part 2 of 7 — Are Two Categorical Variables Related?


Topics in This Part

Section
📐 When to Use It
📊 Two-Way Tables & Expected Counts
🧮 The Test Statistic
📝 Full Worked Example

🔑 Key Concept: The chi-square test for independence uses data from one sample to determine whether two categorical variables are associated (related) or independent.


When to Use the Test for Independence

FeatureDetail
Data sourceOne sample or one group of subjects
VariablesTwo categorical variables measured on each subject
H0H_0The two variables are independent (not associated)
HaH_aThe two variables are associated (dependent)

Example: Survey 300 students and record both grade level (freshman, sophomore, junior, senior) and preferred lunch (pizza, salad, sandwich). Is there an association between grade and lunch preference?


Expected Counts

For each cell in a two-way table:

E=row total×column totalgrand total\boxed{E = \frac{\text{row total} \times \text{column total}}{\text{grand total}}}

This gives the count you would expect if the two variables were truly independent.


Worked Example

Data: A random sample of 200 adults:

FavorOpposeTotal
Male6040100
Female4555100
Total10595200

H0H_0: Gender and opinion are independent.
HaH_a: Gender and opinion are associated.

Expected counts:

FavorOppose
Male100×105200=52.5\frac{100 \times 105}{200} = 52.5100×95200=47.5\frac{100 \times 95}{200} = 47.5
Female100×105200=52.5\frac{100 \times 105}{200} = 52.5100×95200=47.5\frac{100 \times 95}{200} = 47.5

χ2\chi^2 calculation:

χ2=(60−52.5)252.5+(40−47.5)247.5+(45−52.5)252.5+(55−47.5)247.5\chi^2 = \frac{(60-52.5)^2}{52.5} + \frac{(40-47.5)^2}{47.5} + \frac{(45-52.5)^2}{52.5} + \frac{(55-47.5)^2}{47.5}

=56.2552.5+56.2547.5+56.2552.5+56.2547.5= \frac{56.25}{52.5} + \frac{56.25}{47.5} + \frac{56.25}{52.5} + \frac{56.25}{47.5}

=1.071+1.184+1.071+1.184=4.510= 1.071 + 1.184 + 1.071 + 1.184 = 4.510

df=(r−1)(c−1)=(2−1)(2−1)=1df = (r-1)(c-1) = (2-1)(2-1) = 1

Using a χ2\chi^2 table: p≈0.034p \approx 0.034

Conclusion: Since p=0.034<0.05p = 0.034 < 0.05, we reject H0H_0. There is convincing evidence of an association between gender and opinion on this issue.

Independence Test Concepts 🎯

Expected Count Practice 🧮

A two-way table has row totals of 80 and 120, column totals of 90 and 110, and a grand total of 200.

1) Expected count for the top-left cell (row 1, column 1)?

2) Expected count for the bottom-right cell (row 2, column 2)?

3) Degrees of freedom for this 2×22 \times 2 table?

Independence Test Decisions 🔍

Exit Quiz — Test for Independence ✅

Part 3: Chi-Square Test for Homogeneity

📊 Chi-Square Test for Homogeneity

Part 3 of 7 — Comparing Distributions Across Populations


Topics in This Part

Section
📐 Independence vs. Homogeneity
📊 Setting Up the Test
🧮 Worked Example
📝 AP Exam Distinction

🔑 Key Concept: The test for homogeneity uses data from two or more independent samples (or treatment groups) to determine whether the distribution of a single categorical variable is the same across all populations.


Independence vs. Homogeneity

FeatureIndependenceHomogeneity
SamplesOne sampleTwo or more independent samples
VariablesTwo categorical variablesOne categorical variable across groups
H0H_0Variables are independentDistribution is the same across populations
CalculationIdentical formula: χ2=∑(O−E)2/E\chi^2 = \sum (O-E)^2/ESame formula
dfdf(r−1)(c−1)(r-1)(c-1)(r−1)(c−1)(r-1)(c-1)

⚠️ AP Exam: The math is identical. The difference is in the hypotheses and context. Read the problem carefully to determine which test is appropriate.


Hypotheses for Homogeneity

H0:The distribution of [variable] is the same for all [groups].H_0: \text{The distribution of [variable] is the same for all [groups].} Ha:The distribution of [variable] is NOT the same for all [groups].H_a: \text{The distribution of [variable] is NOT the same for all [groups].}


Worked Example

Problem: Two schools were surveyed about favorite subject. Is the distribution of preferences the same?

MathEnglishScienceTotal
School A453025100
School B354025100
Total807050200

H0H_0: The distribution of favorite subject is the same for School A and School B.

Expected counts: (each row total = 100, grand total = 200)

MathEnglishScience
School A100×80200=40\frac{100 \times 80}{200} = 40100×70200=35\frac{100 \times 70}{200} = 35100×50200=25\frac{100 \times 50}{200} = 25
School B404035352525

χ2\chi^2:

χ2=(45−40)240+(30−35)235+(25−25)225+(35−40)240+(40−35)235+(25−25)225\chi^2 = \frac{(45-40)^2}{40} + \frac{(30-35)^2}{35} + \frac{(25-25)^2}{25} + \frac{(35-40)^2}{40} + \frac{(40-35)^2}{35} + \frac{(25-25)^2}{25}

=0.625+0.714+0+0.625+0.714+0=2.678= 0.625 + 0.714 + 0 + 0.625 + 0.714 + 0 = 2.678

df=(2−1)(3−1)=2df = (2-1)(3-1) = 2

Using a χ2\chi^2 table with df=2df = 2: p≈0.262p \approx 0.262

Conclusion: Since p=0.262>0.05p = 0.262 > 0.05, we fail to reject H0H_0. There is not convincing evidence that the distribution of favorite subject differs between the two schools.

Homogeneity Concepts 🎯

Independence or Homogeneity? 🔍

Identify the correct test for each scenario.

Homogeneity Calculation 🧮

Three brands of cereal are compared on sugar level (Low, Medium, High). Samples: Brand A: 50, Brand B: 60, Brand C: 40. Grand total: 150. Column totals: Low = 60, Medium = 50, High = 40.

1) Expected count for Brand A, Low sugar?

2) Expected count for Brand C, High sugar?

3) dfdf for this 3×33 \times 3 table?

Exit Quiz — Test for Homogeneity ✅

Part 4: Conditions and Degrees of Freedom

📊 Conditions and Degrees of Freedom

Part 4 of 7 — When Can You Use the χ2\chi^2 Test?


Topics in This Part

Section
✅ Three Conditions
📐 Degrees of Freedom for Each Test
⚠️ What to Do When Conditions Fail
📝 AP Exam Condition-Checking

🔑 Key Concept: All three chi-square tests (GoF, Independence, Homogeneity) require the same three conditions: Random, 10%, and Large Counts.


The Three Conditions

ConditionRequirementAP Language
RandomData from random sample or randomized experiment"The problem states..." or "We are told..."
10%n<10%n < 10\% of the population (if sampling without replacement)"nn is less than 10% of all [population]"
Large CountsAll expected counts ≥5\geq 5"All expected counts are at least 5" ✓

⚠️ Critical AP Detail: For chi-square, the Large Counts condition uses expected counts, NOT observed counts. This is different from the Large Counts condition for proportions (np^≥10n\hat{p} \geq 10).


Degrees of Freedom Summary

Testdfdf FormulaExample
Goodness of Fitk−1k - 1 (where kk = number of categories)6 sides of a die → df=5df = 5
Independence(r−1)(c−1)(r-1)(c-1)3×43 \times 4 table → df=6df = 6
Homogeneity(r−1)(c−1)(r-1)(c-1)2×32 \times 3 table → df=2df = 2

Why Degrees of Freedom Matter

The χ2\chi^2 distribution changes shape with dfdf:

dfdfShape
11Strongly right-skewed
55Moderately right-skewed
15+15+More symmetric

Higher dfdf shifts the distribution to the right and increases the mean (μ=df\mu = df).


What If Conditions Fail?

ConditionIf It Fails
RandomResults may not generalize — state the limitation
10%Standard errors may be wrong — results are questionable
Large CountsCombine categories or use Fisher exact test (not on AP exam)

🔑 AP Tip: On the AP exam, if an expected count is below 5, you should note this but may still be asked to proceed with the test. State the concern and continue.

Conditions & df Concepts 🎯

Degrees of Freedom Practice 🧮

1) GoF test with 8 categories: df=df =

2) Independence test with a 5×35 \times 3 table: df=df =

3) Homogeneity test comparing 4 groups on a categorical variable with 3 levels: Table is 4×34 \times 3. df=df =

Condition Checking 🔍

Exit Quiz — Conditions & Degrees of Freedom ✅

Part 5: Interpreting Results

📊 Interpreting Results

Part 5 of 7 — Reading χ2\chi^2 Output and Drawing Conclusions


Topics in This Part

Section
📐 Interpreting the χ2\chi^2 Statistic
📊 Using the χ2\chi^2 Table
📝 Writing AP Conclusions
🔍 Follow-Up Analysis

🔑 Key Concept: A large χ2\chi^2 value means the observed data differ substantially from what is expected under H0H_0. The p-value tells you how surprising your χ2\chi^2 would be if H0H_0 were true.


What the χ2\chi^2 Value Tells You

χ2\chi^2 ValueInterpretation
Near 0Observed counts closely match expected → little evidence against H0H_0
ModerateSome discrepancy → may or may not be significant
LargeBig differences → strong evidence against H0H_0

The p-value makes this precise: it gives the probability of getting a χ2\chi^2 as large or larger than yours, assuming H0H_0 is true.


Reading the χ2\chi^2 Table

The table gives right-tail areas for the χ2\chi^2 distribution:

dfdfα=0.10\alpha = 0.10α=0.05\alpha = 0.05α=0.025\alpha = 0.025α=0.01\alpha = 0.01
12.7063.8415.0246.635
24.6055.9917.3789.210
36.2517.8159.34811.345
47.7799.48811.14313.277
59.23611.07012.83315.086

How to use: If χ2=8.5\chi^2 = 8.5 with df=3df = 3: 7.815<8.5<9.3487.815 < 8.5 < 9.348, so 0.025<p<0.050.025 < p < 0.05.


AP Conclusion Template

If p≤αp \leq \alpha: "Since the p-value (p=valuep = \text{value}) is less than α=0.05\alpha = 0.05, we reject H0H_0. There is convincing evidence that [state HaH_a in context]."

If p>αp > \alpha: "Since the p-value (p=valuep = \text{value}) is greater than α=0.05\alpha = 0.05, we fail to reject H0H_0. There is not convincing evidence that [state HaH_a in context]."

⚠️ Never say "accept H0H_0" — say "fail to reject H0H_0."


Follow-Up: Which Cells Drive the Result?

After rejecting H0H_0, examine individual cell contributions (Oi−Ei)2/Ei(O_i - E_i)^2/E_i:

ContributionInterpretation
LargeThis category/cell is a major source of the discrepancy
SmallThis category/cell fits the model well

Also note the direction: Is O>EO > E (more than expected) or O<EO < E (fewer than expected)?

Interpretation Concepts 🎯

Using the χ2\chi^2 Table 🧮

Use the partial table above.

1) χ2=6.0\chi^2 = 6.0, df=2df = 2. Is pp less than or greater than 0.05? Enter "less" or "greater".

2) χ2=10.5\chi^2 = 10.5, df=5df = 5. The p-value is between which two table values? Enter the larger α\alpha boundary (e.g., "0.10").

3) For df=1df = 1, what χ2\chi^2 value gives p=0.05p = 0.05?

Conclusion Writing 🔍

Exit Quiz — Interpreting Results ✅

Part 6: Problem-Solving Workshop

📊 Problem-Solving Workshop

Part 6 of 7 — Full AP Free-Response Practice


Topics in This Part

Section
📝 GoF Worked Example
📝 Independence Worked Example
⚠️ Common AP Mistakes

🔑 Key Concept: Chi-square FRQs follow the same 4-step framework: Hypotheses, Conditions, Calculate, Conclude. Practice writing each step clearly.


Worked Example 1: Goodness-of-Fit

Problem: A company claims its candy mix is 30% red, 20% blue, 20% green, 15% yellow, 15% orange. A random sample of 200 candies yields:

ColorRedBlueGreenYellowOrange
Observed7535322830
Expected6040403030

Step 1 — Hypotheses: H0H_0: The distribution of colors matches the company claim. HaH_a: The distribution of colors does not match the company claim.

Step 2 — Conditions:

  • Random: Random sample stated ✓
  • 10%: 200<10%200 < 10\% of all candies produced ✓
  • Large Counts: All expected counts ≥5\geq 5 (smallest is 30) ✓

Step 3 — Calculate: χ2=(75−60)260+(35−40)240+(32−40)240+(28−30)230+(30−30)230\chi^2 = \frac{(75-60)^2}{60} + \frac{(35-40)^2}{40} + \frac{(32-40)^2}{40} + \frac{(28-30)^2}{30} + \frac{(30-30)^2}{30}

=22560+2540+6440+430+0=3.75+0.625+1.60+0.133+0= \frac{225}{60} + \frac{25}{40} + \frac{64}{40} + \frac{4}{30} + 0 = 3.75 + 0.625 + 1.60 + 0.133 + 0

χ2=6.108,df=5−1=4\chi^2 = 6.108, \quad df = 5-1 = 4

From the table: pp is between 0.100.10 and 0.250.25 (since 6.108<7.7796.108 < 7.779).

Step 4 — Conclude: Since the p-value is greater than α=0.05\alpha = 0.05, we fail to reject H0H_0. There is not convincing evidence that the distribution of candy colors differs from the company claim.

🔍 Follow-up: The red category had the largest contribution (3.75), suggesting there may be more red candies than claimed.


Worked Example 2: Test for Independence

Problem: A random sample of 400 adults records education level and exercise frequency:

≤ 3 days/week> 3 days/weekTotal
No degree12080200
College degree70130200
Total190210400

Step 1 — Hypotheses: H0H_0: Education level and exercise frequency are independent. HaH_a: Education level and exercise frequency are associated.

Step 2 — Conditions:

  • Random: Random sample stated ✓
  • 10%: 400<10%400 < 10\% of all adults ✓
  • Large Counts: Expected counts: E11=200×190400=95E_{11} = \frac{200 \times 190}{400} = 95, E12=105E_{12} = 105, E21=95E_{21} = 95, E22=105E_{22} = 105. All ≥5\geq 5 ✓

Step 3 — Calculate: χ2=(120−95)295+(80−105)2105+(70−95)295+(130−105)2105\chi^2 = \frac{(120-95)^2}{95} + \frac{(80-105)^2}{105} + \frac{(70-95)^2}{95} + \frac{(130-105)^2}{105}

=62595+625105+62595+625105=6.579+5.952+6.579+5.952=25.06= \frac{625}{95} + \frac{625}{105} + \frac{625}{95} + \frac{625}{105} = 6.579 + 5.952 + 6.579 + 5.952 = 25.06

df=(2−1)(2−1)=1df = (2-1)(2-1) = 1. From the table: χ2=25.06>6.635\chi^2 = 25.06 > 6.635 so p<0.01p < 0.01.

Step 4 — Conclude: Since the p-value is less than α=0.05\alpha = 0.05 (in fact less than 0.01), we reject H0H_0. There is convincing evidence of an association between education level and exercise frequency.

⚠️ Common AP Mistakes on Chi-Square FRQs

MistakeFix
Using observed counts for the Large Counts checkMust use expected counts
Forgetting dfdfAlways state dfdf with the formula
Not stating hypotheses in context"H0H_0: Color distribution matches claim" not just "H0H_0: fit"
Saying "accept H0H_0"Say "fail to reject H0H_0"
Confusing independence, homogeneity, and GoFRead the study design carefully
Not showing the χ2\chi^2 formula with substitutionShow ∑(O−E)2/E\sum (O-E)^2/E with at least some terms
Claiming causation from an independence testAssociation ≠ causation (unless randomized experiment)

Workshop Concept Check 🎯

Quick Calculations 🧮

A GoF test: categories A, B, C with observed = 30, 25, 45 and expected = 33.3, 33.3, 33.3 (total = 100, equal proportions).

1) Contribution from category A: (O−E)2/E(O-E)^2/E (round to 2 decimal places)

2) Contribution from category C: (O−E)2/E(O-E)^2/E (round to 2 decimal places)

3) dfdf for this test?

Which Test? 🔍

Exit Quiz — Problem-Solving Workshop ✅

Part 7: Review & Applications

📊 Review & Applications

Part 7 of 7 — Comprehensive Chi-Square Review


Topics in This Part

Section
📋 All Three Tests Side-by-Side
📐 Formula & Condition Summary
📝 Mixed Practice

🔑 Key Concept: This review covers all three chi-square tests. Know when to use each, how to check conditions, and how to write full AP-quality solutions.


Three Chi-Square Tests Compared

FeatureGoodness of FitIndependenceHomogeneity
SamplesOneOneTwo or more
VariablesOne categoricalTwo categoricalOne categorical
H0H_0Distribution matches modelVariables are independentSame distribution across groups
Data FormatOne-way tableTwo-way tableTwo-way table
dfdfk−1k - 1(r−1)(c−1)(r-1)(c-1)(r−1)(c−1)(r-1)(c-1)

Universal Formula

χ2=∑(Oi−Ei)2Ei\chi^2 = \sum \frac{(O_i - E_i)^2}{E_i}

Universal Conditions

ConditionRequirement
RandomRandom sample or randomized experiment
10%n<10%n < 10\% of population
Large CountsAll expected counts ≥5\geq 5

Expected Counts

TestHow to Calculate EE
GoFEi=n×piE_i = n \times p_i (hypothesized proportion)
Independence/HomogeneityE=row total×column totalgrand totalE = \frac{\text{row total} \times \text{column total}}{\text{grand total}}

Decision Guide: Which Test?

QuestionAnswer
Does the data fit a specific model?GoF
Are two variables related (one sample)?Independence
Same distribution across groups (multiple samples)?Homogeneity

Key AP Reminders

  • χ2\chi^2 is always ≥0\geq 0 and always right-tailed
  • Large Counts uses expected counts, not observed
  • Never say "accept H0H_0"
  • Association ≠ causation (unless randomized experiment)
  • Show the formula with substitution on FRQs
  • State dfdf explicitly

Comprehensive Concept Check 🎯

Mixed Practice 🧮

1) GoF test, 4 categories, n=100n = 100, equal proportions. Expected count per category?

2) Independence test, 3×53 \times 5 table. df=df =

3) χ2=7.5\chi^2 = 7.5, df=2df = 2. From the table (α=0.05\alpha = 0.05: 5.991; α=0.025\alpha = 0.025: 7.378). Is pp less than 0.025? Enter "yes" or "no".

Quick Decisions 🔍

Final Exam — Chi-Square Unit ✅