8.4: Comparing Two Means (Independent Samples)
- Page ID
- 58928
\( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)
\( \newcommand{\dsum}{\displaystyle\sum\limits} \)
\( \newcommand{\dint}{\displaystyle\int\limits} \)
\( \newcommand{\dlim}{\displaystyle\lim\limits} \)
\( \newcommand{\id}{\mathrm{id}}\) \( \newcommand{\Span}{\mathrm{span}}\)
( \newcommand{\kernel}{\mathrm{null}\,}\) \( \newcommand{\range}{\mathrm{range}\,}\)
\( \newcommand{\RealPart}{\mathrm{Re}}\) \( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)
\( \newcommand{\Argument}{\mathrm{Arg}}\) \( \newcommand{\norm}[1]{\| #1 \|}\)
\( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)
\( \newcommand{\Span}{\mathrm{span}}\)
\( \newcommand{\id}{\mathrm{id}}\)
\( \newcommand{\Span}{\mathrm{span}}\)
\( \newcommand{\kernel}{\mathrm{null}\,}\)
\( \newcommand{\range}{\mathrm{range}\,}\)
\( \newcommand{\RealPart}{\mathrm{Re}}\)
\( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)
\( \newcommand{\Argument}{\mathrm{Arg}}\)
\( \newcommand{\norm}[1]{\| #1 \|}\)
\( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)
\( \newcommand{\Span}{\mathrm{span}}\) \( \newcommand{\AA}{\unicode[.8,0]{x212B}}\)
\( \newcommand{\vectorA}[1]{\vec{#1}} % arrow\)
\( \newcommand{\vectorAt}[1]{\vec{\text{#1}}} % arrow\)
\( \newcommand{\vectorB}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\( \newcommand{\vectorC}[1]{\textbf{#1}} \)
\( \newcommand{\vectorD}[1]{\overrightarrow{#1}} \)
\( \newcommand{\vectorDt}[1]{\overrightarrow{\text{#1}}} \)
\( \newcommand{\vectE}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash{\mathbf {#1}}}} \)
\( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\(\newcommand{\longvect}{\overrightarrow}\)
\( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)
\(\newcommand{\avec}{\mathbf a}\) \(\newcommand{\bvec}{\mathbf b}\) \(\newcommand{\cvec}{\mathbf c}\) \(\newcommand{\dvec}{\mathbf d}\) \(\newcommand{\dtil}{\widetilde{\mathbf d}}\) \(\newcommand{\evec}{\mathbf e}\) \(\newcommand{\fvec}{\mathbf f}\) \(\newcommand{\nvec}{\mathbf n}\) \(\newcommand{\pvec}{\mathbf p}\) \(\newcommand{\qvec}{\mathbf q}\) \(\newcommand{\svec}{\mathbf s}\) \(\newcommand{\tvec}{\mathbf t}\) \(\newcommand{\uvec}{\mathbf u}\) \(\newcommand{\vvec}{\mathbf v}\) \(\newcommand{\wvec}{\mathbf w}\) \(\newcommand{\xvec}{\mathbf x}\) \(\newcommand{\yvec}{\mathbf y}\) \(\newcommand{\zvec}{\mathbf z}\) \(\newcommand{\rvec}{\mathbf r}\) \(\newcommand{\mvec}{\mathbf m}\) \(\newcommand{\zerovec}{\mathbf 0}\) \(\newcommand{\onevec}{\mathbf 1}\) \(\newcommand{\real}{\mathbb R}\) \(\newcommand{\twovec}[2]{\left[\begin{array}{r}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\ctwovec}[2]{\left[\begin{array}{c}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\threevec}[3]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\cthreevec}[3]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\fourvec}[4]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\cfourvec}[4]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\fivevec}[5]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\cfivevec}[5]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\mattwo}[4]{\left[\begin{array}{rr}#1 \amp #2 \\ #3 \amp #4 \\ \end{array}\right]}\) \(\newcommand{\laspan}[1]{\text{Span}\{#1\}}\) \(\newcommand{\bcal}{\cal B}\) \(\newcommand{\ccal}{\cal C}\) \(\newcommand{\scal}{\cal S}\) \(\newcommand{\wcal}{\cal W}\) \(\newcommand{\ecal}{\cal E}\) \(\newcommand{\coords}[2]{\left\{#1\right\}_{#2}}\) \(\newcommand{\gray}[1]{\color{gray}{#1}}\) \(\newcommand{\lgray}[1]{\color{lightgray}{#1}}\) \(\newcommand{\rank}{\operatorname{rank}}\) \(\newcommand{\row}{\text{Row}}\) \(\newcommand{\col}{\text{Col}}\) \(\renewcommand{\row}{\text{Row}}\) \(\newcommand{\nul}{\text{Nul}}\) \(\newcommand{\var}{\text{Var}}\) \(\newcommand{\corr}{\text{corr}}\) \(\newcommand{\len}[1]{\left|#1\right|}\) \(\newcommand{\bbar}{\overline{\bvec}}\) \(\newcommand{\bhat}{\widehat{\bvec}}\) \(\newcommand{\bperp}{\bvec^\perp}\) \(\newcommand{\xhat}{\widehat{\xvec}}\) \(\newcommand{\vhat}{\widehat{\vvec}}\) \(\newcommand{\uhat}{\widehat{\uvec}}\) \(\newcommand{\what}{\widehat{\wvec}}\) \(\newcommand{\Sighat}{\widehat{\Sigma}}\) \(\newcommand{\lt}{<}\) \(\newcommand{\gt}{>}\) \(\newcommand{\amp}{&}\) \(\definecolor{fillinmathshade}{gray}{0.9}\)Sometimes a single sample isn't enough. Instead, we want to compare two groups — to test whether one mean is higher (or lower) than the other.
In these situations, we'll use a two-sample t-test. If we do not assume equal variances between the groups, we'll specifically use Welch's t-test.
Questions That Involve Comparing Two Means
Here are five real-world questions where comparing group averages is a natural fit:
- Do students who attend tutoring score higher on exams than those who don't?
- Is it colder on average in Leadville, Colorado than in Anchorage, Alaska?
- Do iPhone users spend more on monthly mobile plans than Android users?
- Does a new fertilizer increase the average crop yield compared to the standard?
- Do influencers post more times per day on TikTok than Twitter users do on Twitter?
All of these questions can be framed as tests comparing the means of two independent samples.
Welch's t-test Statistic
Used when the sample sizes or variances may be unequal:
\[ T = \frac{\bar{x}_1 - \bar{x}_2}{\sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}} \]
We'll use this statistic to test differences between two groups. This test statistic follows the Student t-distribution, just like the t-tests we saw in the last section.
Conditions for Using This Test
Before running a two-sample t-test, verify that these conditions are met:
- Independence: The two samples are independent of each other, and observations within each sample are independent.
- Randomness: Both samples were collected using a random process.
- Approximate normality: Each group's data should be roughly symmetric, or the sample sizes should be large enough (generally \( n \geq 30 \) per group) for the Central Limit Theorem to apply.
Let's try it out in a real example.
Example: "Is it colder in Leadville than Anchorage?"
You're arguing with a friend: They claim it's obviously colder in Alaska. But you're not so sure — Leadville, Colorado is at 10,000 feet after all.
You decide to test whether the average winter temperature in Leadville is colder than in Anchorage.
Step 1: State the Hypotheses
You'd like to test:
- \( H_0: \mu_{Leadville} = \mu_{Anchorage} \)
- \( H_A: \mu_{Leadville} < \mu_{Anchorage} \)
This is a left-tailed test. You're looking for evidence that winter temperatures are lower in Leadville.
Step 2: Calculate the Appropriate Test Statistic
You collect 10 average January temperatures from each city (in °F):
- Leadville: \( \bar{x}_1 = 12.2, s_1 = 4.8, n_1 = 10 \)
- Anchorage: \( \bar{x}_2 = 17.6, s_2 = 6.2, n_2 = 10 \)
This is a two-sample independent t-test. Sample standard deviations are not equal → We use Welch's t-test.
Plug into the equation:
\[ T = \frac{12.2 - 17.6}{\sqrt{\frac{(4.8)^2}{10} + \frac{(6.2)^2}{10}}} = \frac{-5.4}{\sqrt{2.304 + 3.844}} = \frac{-5.4}{\sqrt{6.148}} \approx \frac{-5.4}{2.479} \approx -2.18 \]
Step 3: Choose a Significance Level
You select the standard \( \alpha = 0.05 \) as your cutoff for a Type I error (false positive).
Step 4: Compute p-value Using Technology
The degrees of freedom here are computed using the Welch–Satterthwaite formula (opens in new window). Generally, this calculation is included in whatever technology you're using. In some cases, we may use a simplified, conservative option where the degrees of freedom is the smaller of \( n_1 - 1 \) and \( n_2 - 1 \), but this does typically result in slightly larger p-values.
Use technology to find \( P(t < -2.18) \). We get \( p \approx 0.023 \).
Step 5: Make Your Conclusion
Since \( 0.023 < 0.05 \), you reject \( H_0 \).
Conclusion: There is statistically significant evidence that winters in Leadville are colder than in Anchorage.
Example 2: "Do students in the new program score differently on the final exam?" (Try It)
A college introduces a new hybrid learning format for its introductory statistics course. An administrator wants to know whether students in the new format score differently on the final exam than students in the traditional format.
A random sample from each group yields the following:
- Hybrid format: \( \bar{x}_1 = 81.4, s_1 = 9.2, n_1 = 22 \)
- Traditional format: \( \bar{x}_2 = 76.8, s_2 = 11.5, n_2 = 25 \)
Test at \( \alpha = 0.05 \) whether there is evidence of a difference in mean scores between the two formats.
- State \( H_0 \) and \( H_A \). What type of test is this — left, right, or two-tailed?
- Calculate the Welch's t test statistic.
- Use technology to find the p-value (use the conservative \( df = \min(n_1 - 1, n_2 - 1) = 21 \) if a calculator isn't available).
- Make a decision and write a conclusion in context.
Hint: Since the question asks whether scores are different (not which format is better), this is a two-tailed test. Remember to double the one-tail p-value.
Tip: The denominator of Welch's statistic combines the variance from both groups separately — do not pool the standard deviations. This is what distinguishes Welch's t-test from the equal-variance (pooled) t-test.
Related Video
Looking Ahead: Comparing Proportions
We've now compared two means — but many research questions ask about proportions rather than averages. Instead of asking "do two groups have different mean scores?", we might ask "do two groups have different rates?" For example:
- Is the graduation rate higher in one program than another?
- Do two age groups differ in their support for a policy?
- Does one treatment have a higher success rate than another?
In the next section, we'll extend our toolkit to handle exactly these questions using a two-proportion z-test. The five-step framework stays the same — only the test statistic and distribution change.


