Skip to content

Inference for Regression

Perform inference for the slope of a regression line using t-tests and confidence intervals.

Written and reviewed by the Study Mondo Education TeamLast updated
🎯⭐ INTERACTIVE LESSON

Try the Interactive Version!

Learn step-by-step with practice exercises built right in.

Start Interactive Lesson →

📈 Inference for the Slope of a Regression Line

Test of Significance for Slope

Null hypothesis: H0:β=0H_0: \beta = 0 (no linear relationship)

Test statistic: t=b−0SE(b)=bSE(b)t = \frac{b - 0}{SE(b)} = \frac{b}{SE(b)}

Where:

  • bb = slope of regression line
  • SE(b)SE(b) = standard error of slope
  • Degrees of freedom: df=n−2df = n - 2

Interpretation:

  • If ∣t∣>t∗|t| > t^*, reject H0H_0: slope is significantly different from 0
  • If we reject, there IS a significant linear relationship

Conditions (LINER)

  1. Linear: Scatterplot shows linear trend
  2. Independent: Observations independent
  3. Normal: Residuals approximately normal (histogram or QQ-plot)
  4. Equal SD: Constant vertical spread (residual plot shows homogeneity)
  5. Random: Random sample or random assignment

Never skip: Always state all five conditions and explain how you checked each.

Worked Example

Data: Study hours vs. Exam score (n = 20 students)

Regression: Score^=65+4.2⋅Hours\hat{\text{Score}} = 65 + 4.2 \cdot \text{Hours}

SE(b)=1.1SE(b) = 1.1, so t=4.21.1=3.82t = \frac{4.2}{1.1} = 3.82

df=18df = 18; t∗=2.101t^* = 2.101 (two-tailed, α=0.05\alpha = 0.05)

Since ∣3.82∣>2.101|3.82| > 2.101, reject H0H_0.

Conclusion: There is significant evidence of a linear relationship between hours studied and exam score.

Confidence Interval for Slope

b±t∗⋅SE(b)b \pm t^* \cdot SE(b)

Example: 4.2±2.101(1.1)=4.2±2.31=(1.89,6.51)4.2 \pm 2.101(1.1) = 4.2 \pm 2.31 = (1.89, 6.51)

Interpretation: We are 95% confident the true slope is between 1.89 and 6.51 points per hour.

Common Mistakes

❌ Not checking LINER conditions (major point deduction) ❌ Using α=0.05\alpha = 0.05 without stating it ❌ Confusing test for slope with correlation significance (similar but different) ❌ Forgetting df=n−2df = n - 2 (not n−1n - 1)

Decision Rule

  • If ∣t∣>tn−2,α/2∗|t| > t^*_{n-2, \alpha/2}, reject H0H_0
  • If p-value <α< \alpha, reject H0H_0

AP Exam Tip

State all five conditions and HOW you checked each (e.g., "Residual plot shows random scatter, supporting linearity"). Name the test: "t-test for the slope." Report the test statistic, degrees of freedom, and p-value (or critical value). Always conclude in context.

📚 Practice Problems

1Problem 1medium

❓ Question:

A regression of study hours (x) on test scores (y) gives slope b₁ = 5.2 with SE = 1.3, n = 20. Construct a 95% confidence interval for the true slope β₁.

💡 Show Solution

Step 1: Identify given information Slope: b₁ = 5.2 Standard error: SE = 1.3 Sample size: n = 20 Confidence level: 95%

Step 2: Find degrees of freedom df = n - 2 = 20 - 2 = 18 (Use n-2 for regression, not n-1)

Step 3: Find t* critical value From t-table with df = 18, 95% confidence: t* = 2.101

Step 4: Calculate margin of error ME = t* × SE ME = 2.101 × 1.3 ME ≈ 2.73

Step 5: Construct confidence interval CI = b₁ ± ME CI = 5.2 ± 2.73 CI = (2.47, 7.93)

Step 6: Interpret "We are 95% confident that for each additional hour studied, the true mean increase in test score is between 2.47 and 7.93 points."

Note: Since 0 is NOT in the interval, there is significant evidence of a positive relationship (can reject H₀: β₁ = 0).

Answer: 95% CI: (2.47, 7.93) points per hour

2Problem 2medium

❓ Question:

Test H₀: β₁ = 0 vs Hₐ: β₁ ≠ 0 given b₁ = 3.5, SE = 1.2, n = 25, α = 0.05.

💡 Show Solution

Step 1: Set up hypotheses H₀: β₁ = 0 (no relationship) Hₐ: β₁ ≠ 0 (relationship exists)

Two-tailed test, α = 0.05

Step 2: Check conditions LINEAR: Assume scatterplot is linear ✓ INDEPENDENT: Assume random sample, n < 10% population ✓ NORMAL: Residuals approximately normal ✓ EQUAL VARIANCE: Residual plot shows constant spread ✓ RANDOM: Random sample ✓

(LINE conditions for regression inference)

Step 3: Calculate test statistic df = n - 2 = 25 - 2 = 23

t = (b₁ - 0)/SE t = 3.5/1.2 t ≈ 2.917

Step 4: Find p-value From t-table with df = 23, two-tailed: t = 2.917 is between t = 2.807 (p = 0.01) and t = 3.767 (p = 0.001)

So: 0.001 < p-value < 0.01

More precisely: p-value ≈ 0.0077

Step 5: Make decision p-value (0.0077) < α (0.05) REJECT H₀

Step 6: Conclusion in context "There is significant evidence (p = 0.008) that a linear relationship exists between x and y. The slope is significantly different from zero."

Answer: t = 2.92, p-value ≈ 0.008. Reject H₀. Significant evidence of linear relationship.

3Problem 3medium

❓ Question:

What are the conditions (LINE) for inference in regression? Explain each briefly.

💡 Show Solution

The LINE conditions for regression inference:

L - LINEAR Relationship between x and y is linear Check: Scatterplot should show linear pattern Residual plot should show no curve

I - INDEPENDENT
Observations are independent Check: Random sampling n < 10% of population (if sampling without replacement) No time series or repeated measures

N - NORMAL Residuals are approximately normally distributed Check: Histogram or normal probability plot of residuals Not critical if n is large (n ≥ 30) Just need no strong skewness or outliers

E - EQUAL VARIANCE (also called homoscedasticity) Variability of y is constant for all x Check: Residual plot shows roughly equal vertical spread No fan shape or other pattern in spread

Why these matter:

  • LINEAR: For model to be appropriate
  • INDEPENDENT: For formulas to be valid
  • NORMAL: For t-distribution to apply (especially small samples)
  • EQUAL VARIANCE: For standard errors to be correct

If violations:

  • Not linear → transform or use nonlinear model
  • Not independent → use different methods (time series, etc.)
  • Not normal → okay if n ≥ 30; otherwise transform
  • Not equal variance → transform or use weighted regression

Answer: LINE = Linear relationship, Independent observations, Normal residuals, Equal variance. Check using scatterplot, residual plot, and normal probability plot.

4Problem 4medium

❓ Question:

Computer output shows: b₁ = 2.4, SE(b₁) = 0.8, t = 3.0, p = 0.006, n = 22. Interpret the p-value in context.

💡 Show Solution

Step 1: Identify the test Testing: H₀: β₁ = 0 (no relationship) Against: Hₐ: β₁ ≠ 0 (relationship exists)

Given: p-value = 0.006

Step 2: What p-value means statistically The probability of observing a slope as extreme as 2.4 (or more extreme) IF the true slope is actually 0.

Step 3: Interpret in context "If there were truly no linear relationship between x and y (β₁ = 0), the probability of obtaining a sample slope of 2.4 or more extreme (in either direction) is 0.006, or 0.6%."

Step 4: Practical interpretation This is very unlikely (less than 1% chance)!

Therefore: Strong evidence AGAINST H₀ The relationship is statistically significant.

Step 5: Decision at α = 0.05 Since p-value (0.006) < α (0.05): REJECT H₀

Conclusion: "There is strong evidence of a significant linear relationship. The slope is significantly different from zero (p = 0.006)."

Step 6: What this does NOT mean ✗ Does not mean slope is definitely 2.4 ✗ Does not mean x causes y ✗ Does not mean model fits well (could still have problems) ✓ Only means: slope significantly different from zero

Answer: If true slope were 0, probability of getting b₁ = 2.4 or more extreme is only 0.006. This provides strong evidence the slope is not zero - there is a significant linear relationship.

5Problem 5hard

❓ Question:

Why do we use t-distribution with df = n-2 for regression inference instead of df = n-1?

💡 Show Solution

Step 1: Compare to one-sample t-test One-sample t-test: df = n - 1

  • Estimate 1 parameter: μ
  • Lose 1 df

Regression: df = n - 2

  • Estimate 2 parameters: β₀ AND β₁
  • Lose 2 df

Step 2: What we're estimating In regression, we estimate:

  1. Intercept (β₀)
  2. Slope (β₁)

Both use up degrees of freedom!

Step 3: Degrees of freedom explained Start with n observations

  • Use one to estimate β₀ (intercept)
  • Use one to estimate β₁ (slope)
  • Left with n - 2 for error estimation

df = n - 2

Step 4: Why it matters Smaller df → wider t* critical values → wider CIs

Example: n = 10, 95% confidence

  • One-sample (df = 9): t* = 2.262
  • Regression (df = 8): t* = 2.306

Regression CI slightly wider (more uncertainty).

Step 5: As n increases For large n, the difference is minimal:

  • df = 100 vs 98 → nearly same t*
  • Both approach z* = 1.96

Step 6: General pattern Degrees of freedom = n - (number of parameters estimated)

  • Mean only: n - 1
  • Regression: n - 2
  • Multiple regression with k predictors: n - k - 1

Answer: We estimate TWO parameters (β₀ and β₁), so we lose 2 degrees of freedom, giving df = n - 2. This accounts for the extra uncertainty from estimating both intercept and slope.

Explain using:

⚠️ Common Mistakes: Inference for Regression

Avoid these 3 frequent errors

❓ Frequently Asked Questions

What is Inference for Regression?▾
Perform inference for the slope of a regression line using t-tests and confidence intervals.
How can I study Inference for Regression effectively?▾
Start by reading the study notes and working through the examples on this page. Then use the flashcards to test your recall. Practice with the 5 problems provided, checking solutions as you go. Regular review and active practice are key to retention.
Is this Inference for Regression study guide free?▾
Yes — all study notes, flashcards, and practice problems for Inference for Regression on Study Mondo are free to access. No account is needed.
What course covers Inference for Regression?▾
Inference for Regression is part of the AP Statistics course on Study Mondo, specifically in the Unit 9: Inference for Quantitative Data — Slopes section. You can explore the full course for more related topics and practice resources.
Are there practice problems for Inference for Regression?▾
Yes, this page includes 5 practice problems with detailed solutions. Each problem includes a step-by-step explanation to help you understand the approach.