8.1: Why hypothesis testing and Type I and Type II Errors
- Page ID
- 58924
\( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)
\( \newcommand{\dsum}{\displaystyle\sum\limits} \)
\( \newcommand{\dint}{\displaystyle\int\limits} \)
\( \newcommand{\dlim}{\displaystyle\lim\limits} \)
\( \newcommand{\id}{\mathrm{id}}\) \( \newcommand{\Span}{\mathrm{span}}\)
( \newcommand{\kernel}{\mathrm{null}\,}\) \( \newcommand{\range}{\mathrm{range}\,}\)
\( \newcommand{\RealPart}{\mathrm{Re}}\) \( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)
\( \newcommand{\Argument}{\mathrm{Arg}}\) \( \newcommand{\norm}[1]{\| #1 \|}\)
\( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)
\( \newcommand{\Span}{\mathrm{span}}\)
\( \newcommand{\id}{\mathrm{id}}\)
\( \newcommand{\Span}{\mathrm{span}}\)
\( \newcommand{\kernel}{\mathrm{null}\,}\)
\( \newcommand{\range}{\mathrm{range}\,}\)
\( \newcommand{\RealPart}{\mathrm{Re}}\)
\( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)
\( \newcommand{\Argument}{\mathrm{Arg}}\)
\( \newcommand{\norm}[1]{\| #1 \|}\)
\( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)
\( \newcommand{\Span}{\mathrm{span}}\) \( \newcommand{\AA}{\unicode[.8,0]{x212B}}\)
\( \newcommand{\vectorA}[1]{\vec{#1}} % arrow\)
\( \newcommand{\vectorAt}[1]{\vec{\text{#1}}} % arrow\)
\( \newcommand{\vectorB}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\( \newcommand{\vectorC}[1]{\textbf{#1}} \)
\( \newcommand{\vectorD}[1]{\overrightarrow{#1}} \)
\( \newcommand{\vectorDt}[1]{\overrightarrow{\text{#1}}} \)
\( \newcommand{\vectE}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash{\mathbf {#1}}}} \)
\( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\(\newcommand{\longvect}{\overrightarrow}\)
\( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)
\(\newcommand{\avec}{\mathbf a}\) \(\newcommand{\bvec}{\mathbf b}\) \(\newcommand{\cvec}{\mathbf c}\) \(\newcommand{\dvec}{\mathbf d}\) \(\newcommand{\dtil}{\widetilde{\mathbf d}}\) \(\newcommand{\evec}{\mathbf e}\) \(\newcommand{\fvec}{\mathbf f}\) \(\newcommand{\nvec}{\mathbf n}\) \(\newcommand{\pvec}{\mathbf p}\) \(\newcommand{\qvec}{\mathbf q}\) \(\newcommand{\svec}{\mathbf s}\) \(\newcommand{\tvec}{\mathbf t}\) \(\newcommand{\uvec}{\mathbf u}\) \(\newcommand{\vvec}{\mathbf v}\) \(\newcommand{\wvec}{\mathbf w}\) \(\newcommand{\xvec}{\mathbf x}\) \(\newcommand{\yvec}{\mathbf y}\) \(\newcommand{\zvec}{\mathbf z}\) \(\newcommand{\rvec}{\mathbf r}\) \(\newcommand{\mvec}{\mathbf m}\) \(\newcommand{\zerovec}{\mathbf 0}\) \(\newcommand{\onevec}{\mathbf 1}\) \(\newcommand{\real}{\mathbb R}\) \(\newcommand{\twovec}[2]{\left[\begin{array}{r}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\ctwovec}[2]{\left[\begin{array}{c}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\threevec}[3]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\cthreevec}[3]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\fourvec}[4]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\cfourvec}[4]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\fivevec}[5]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\cfivevec}[5]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\mattwo}[4]{\left[\begin{array}{rr}#1 \amp #2 \\ #3 \amp #4 \\ \end{array}\right]}\) \(\newcommand{\laspan}[1]{\text{Span}\{#1\}}\) \(\newcommand{\bcal}{\cal B}\) \(\newcommand{\ccal}{\cal C}\) \(\newcommand{\scal}{\cal S}\) \(\newcommand{\wcal}{\cal W}\) \(\newcommand{\ecal}{\cal E}\) \(\newcommand{\coords}[2]{\left\{#1\right\}_{#2}}\) \(\newcommand{\gray}[1]{\color{gray}{#1}}\) \(\newcommand{\lgray}[1]{\color{lightgray}{#1}}\) \(\newcommand{\rank}{\operatorname{rank}}\) \(\newcommand{\row}{\text{Row}}\) \(\newcommand{\col}{\text{Col}}\) \(\renewcommand{\row}{\text{Row}}\) \(\newcommand{\nul}{\text{Nul}}\) \(\newcommand{\var}{\text{Var}}\) \(\newcommand{\corr}{\text{corr}}\) \(\newcommand{\len}[1]{\left|#1\right|}\) \(\newcommand{\bbar}{\overline{\bvec}}\) \(\newcommand{\bhat}{\widehat{\bvec}}\) \(\newcommand{\bperp}{\bvec^\perp}\) \(\newcommand{\xhat}{\widehat{\xvec}}\) \(\newcommand{\vhat}{\widehat{\vvec}}\) \(\newcommand{\uhat}{\widehat{\uvec}}\) \(\newcommand{\what}{\widehat{\wvec}}\) \(\newcommand{\Sighat}{\widehat{\Sigma}}\) \(\newcommand{\lt}{<}\) \(\newcommand{\gt}{>}\) \(\newcommand{\amp}{&}\) \(\definecolor{fillinmathshade}{gray}{0.9}\)Before we dive into numbers or formulas, we need to zoom out and ask an essential question:
In hypothesis testing, we start with a default or preexisting belief or assumption about the world and then use data to test whether that belief is reasonable or needs to be reconsidered.
Example Scenarios: What Are We Testing?
Let’s look at four common scenarios where hypothesis testing appears in real life. Each one sets up a basic question we're trying to answer and helps illustrate what a hypothesis really is. In each scenario, we need to have a default assumption that mimics the phrase "innocent until proven guilty".
- Is this new drug more effective than the standard treatment?
This question is not about a single patient; it’s about a larger population. We’re comparing two treatments, looking for evidence that one is better than the other, and verifying that any observed difference is unlikely to be due to chance. Our default assumption, playing it as safe as possible, is that a new drug is not any more effective than the old treatment until demonstrated to be.
- Does an after-school program improve math scores?
In this scenario we are comparing the average math scores between students who attend an after-school program and those who don't. Our default hypothesis is that the scores should be the same. If the data supports the alternative idea that the after-school program helps, we see how likely it is that the data is meaningful or a fluke.
- Is this machine producing parts that meet design standards?
We have a certain threshold for the proportion of parts that are in tolerance. If we are at or below that proportion, we don't need to do anything and that would be our default. After sampling parts, we can see if the data indicates that too many parts are out of spec.
- Will a local law pass?
If more than 50% of voters choose to support a law in an election it will pass. Suppose we conduct a political poll and find that the majority of voters in our poll do support the bill. How likely is it that we chose a representative sample and when the actual election arrives we will see the same trend?
In all of these examples, we are doing some form of comparison: new vs. old, program vs. no program, actual vs. expected, poll vs. election. And in each case, we are trying to use evidence (our sample) to make a judgment about the entire population. A hypothesis gives us something to test.
Definition: Null Hypothesis
The null hypothesis, written as \( H_0 \), is a formal statement that there is no effect, no difference, or no change in the population. It represents a starting assumption that we test using data. Without knowing data, we need to come into a scenario without any bias, so the null hypothesis reflects that by often assuming the least.
In a hypothesis test, we conjecture that the null hypothesis is true, and then use sample data to evaluate whether there is strong enough evidence to reject it. If our sample statistics are very far away from what the null hypothesis predicts, it lends evidence that the null hypothesis, in fact, is not correct.
The reason we do this, is that if our data suggests the null hypothesis is false, but not very strongly, then there was a chance we arrived at that data due to a non-representative sample. We are effectively determining if the data we collect supports a hypothesis due to a sampling fluke.
We state our null hypothesis using a mathematical formula as shown below. Note that some of the hypotheses are "two-sided", or assume equality, and some hypotheses are one-sided and assume an inequality. There are certain fields where only two-sided hypotheses are used.
Examples of Null Hypotheses:
- The new drug has the same recovery rate as the current one. In this scenario we don't want to assume that the new drug is any better or worse than the old treatment until we test it. We could compare the average recovery rates between the drugs. In this case, \[ H_0 : \mu_{\text{new}} = \mu_{\text{old}}\]
- The average test score from math students before an after school program is 0.68. In this scenario we are leaning on old information. The simplest assumption is that student performance has not improved or gotten worse due to the after school support program. Our null hypothesis can be written: \[ H_0 : \mu_{\text{support}\} = 0.68\]
- Our company decides to shut down part production if more than 2% of parts are out of tolerance. We sample parts to find a proportion of out of spec parts \( p_1\). Our null would take the form \[ H_0 : p_1 \leq 0.02\]
- A bill fails to pass if less than 50% of voters vote for it. We do an opinion poll before the election to try to guess if it will pass. We have a null hypothesis that it will not pass that we investigate: \[ H_0 : p_{\text{poll}} \leq 0.5 \]
Dealing With Rejection
Once we have a null hypothesis, we will only do one of two things with it. If our evidence against the hypothesis is strong enough, we may reject the null hypothesis in favor of an alternative. If our evidence is not strong, we fail to reject the null hypothesis. This doesn't mean that we think it is necessarily true or false, we just don't have good enough evidence to back up a claim of rejection. Depending on context, the data may still imply the null is false, just without being statistically significant.

The Two Types of Errors
Our conclusion might not match what's actually true in the population.
We can think of the population parameters, the true values, as existing but not known. Assuming correct sampling, our statistics should be relatively close to the true population values, but that isn't guaranteed. This is a statement of the Central Limit Theorem! For example, the null hypothesis could be true, but we could get unlucky with a random sample and calculate sample statistics far enough away to warrant rejection of the null. In this case, we accidentally reject a true hypothesis.
Another possibility is that the null hypothesis is wrong, but we end up with sample data that is coincidentally close to the null parameters. In this case we decide not to reject null hypothesis when we should have. These two possibilities are summarized below:
Type I Error
This happens when the null hypothesis is actually true but our test leads us to wrongly reject it.
“We thought there was an effect, but there wasn’t.”
Type II Error
This happens when the null hypothesis is actually false but the test fails to detect the difference, and we wrongly keep it.
“There really was a difference, but we didn’t find it.”
Error Decision Table
This table shows all four possible situations:
| Reality (Truth) | We Reject Null Hypothesis | We Fail to Reject Null Hypothesis |
|---|---|---|
| Null is True | ❌ Type I Error | ✅ Correct Decision |
| Null is False | ✅ Correct Decision | ❌ Type II Error |
It is important to remember that statistics will never “prove” anything with absolute certainty, but it can help us make informed decisions and assess how much risk of error we’re willing to tolerate.
Alpha and Beta
We have two values that we use to describe the probability of comitting a Type I or Type II error:
Note that Type I errors can only occur if the null hypothesis is true. Furthermore, since we are the ones deciding to reject or not, we can structure our test to control for a certain probability of a Type I error occurring. This probability of a Type I error is called the significance level and denoted below:
\[ \alpha = P(\text{Committing a Type I error } | \text{ Null hypothesis is true})\]
We determine or are assigned an \(\alpha\) that we will use for the hypothesis test before analysis is done. Our decision to reject is based on the value we chose, which will inform the rejection. For example, a common value is \(\alpha = 0.05\). This is equivalent to being comfortable with the idea that if the null hypothesis is true, we would still reject it in 5% of possible samples by accident.
Note that \(\alpha \) only represents a probability if the null is true. If the null is false, we cannot commit a Type I error.
The measure of probability of a Type II error is defined as:
\[ \beta = P(\text{Committing a Type II error } | \text{ Null hypothesis is false})\]
In general we do not know \(\beta\) but we can estimate it. We define the power of the test = \(1 - \beta\) to describe the likelihood of correctly rejecting the null.
In the process of hypothesis testing, we predetermine the significance level, but do not actually know if the null is true or not. We make our decision based on how comfortable we are committing a Type I error. However, if the null is in fact false, too low of an \(\alpha\) will increase \(\beta\) as we are too hesitant to reject.


