12.4: Rejection Regions and Risk
- Page ID
- 64769
\( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)
\( \newcommand{\dsum}{\displaystyle\sum\limits} \)
\( \newcommand{\dint}{\displaystyle\int\limits} \)
\( \newcommand{\dlim}{\displaystyle\lim\limits} \)
\( \newcommand{\id}{\mathrm{id}}\) \( \newcommand{\Span}{\mathrm{span}}\)
( \newcommand{\kernel}{\mathrm{null}\,}\) \( \newcommand{\range}{\mathrm{range}\,}\)
\( \newcommand{\RealPart}{\mathrm{Re}}\) \( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)
\( \newcommand{\Argument}{\mathrm{Arg}}\) \( \newcommand{\norm}[1]{\| #1 \|}\)
\( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)
\( \newcommand{\Span}{\mathrm{span}}\)
\( \newcommand{\id}{\mathrm{id}}\)
\( \newcommand{\Span}{\mathrm{span}}\)
\( \newcommand{\kernel}{\mathrm{null}\,}\)
\( \newcommand{\range}{\mathrm{range}\,}\)
\( \newcommand{\RealPart}{\mathrm{Re}}\)
\( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)
\( \newcommand{\Argument}{\mathrm{Arg}}\)
\( \newcommand{\norm}[1]{\| #1 \|}\)
\( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)
\( \newcommand{\Span}{\mathrm{span}}\) \( \newcommand{\AA}{\unicode[.8,0]{x212B}}\)
\( \newcommand{\vectorA}[1]{\vec{#1}} % arrow\)
\( \newcommand{\vectorAt}[1]{\vec{\text{#1}}} % arrow\)
\( \newcommand{\vectorB}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\( \newcommand{\vectorC}[1]{\textbf{#1}} \)
\( \newcommand{\vectorD}[1]{\overrightarrow{#1}} \)
\( \newcommand{\vectorDt}[1]{\overrightarrow{\text{#1}}} \)
\( \newcommand{\vectE}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash{\mathbf {#1}}}} \)
\( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\(\newcommand{\longvect}{\overrightarrow}\)
\( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)
\(\newcommand{\avec}{\mathbf a}\) \(\newcommand{\bvec}{\mathbf b}\) \(\newcommand{\cvec}{\mathbf c}\) \(\newcommand{\dvec}{\mathbf d}\) \(\newcommand{\dtil}{\widetilde{\mathbf d}}\) \(\newcommand{\evec}{\mathbf e}\) \(\newcommand{\fvec}{\mathbf f}\) \(\newcommand{\nvec}{\mathbf n}\) \(\newcommand{\pvec}{\mathbf p}\) \(\newcommand{\qvec}{\mathbf q}\) \(\newcommand{\svec}{\mathbf s}\) \(\newcommand{\tvec}{\mathbf t}\) \(\newcommand{\uvec}{\mathbf u}\) \(\newcommand{\vvec}{\mathbf v}\) \(\newcommand{\wvec}{\mathbf w}\) \(\newcommand{\xvec}{\mathbf x}\) \(\newcommand{\yvec}{\mathbf y}\) \(\newcommand{\zvec}{\mathbf z}\) \(\newcommand{\rvec}{\mathbf r}\) \(\newcommand{\mvec}{\mathbf m}\) \(\newcommand{\zerovec}{\mathbf 0}\) \(\newcommand{\onevec}{\mathbf 1}\) \(\newcommand{\real}{\mathbb R}\) \(\newcommand{\twovec}[2]{\left[\begin{array}{r}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\ctwovec}[2]{\left[\begin{array}{c}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\threevec}[3]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\cthreevec}[3]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\fourvec}[4]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\cfourvec}[4]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\fivevec}[5]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\cfivevec}[5]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\mattwo}[4]{\left[\begin{array}{rr}#1 \amp #2 \\ #3 \amp #4 \\ \end{array}\right]}\) \(\newcommand{\laspan}[1]{\text{Span}\{#1\}}\) \(\newcommand{\bcal}{\cal B}\) \(\newcommand{\ccal}{\cal C}\) \(\newcommand{\scal}{\cal S}\) \(\newcommand{\wcal}{\cal W}\) \(\newcommand{\ecal}{\cal E}\) \(\newcommand{\coords}[2]{\left\{#1\right\}_{#2}}\) \(\newcommand{\gray}[1]{\color{gray}{#1}}\) \(\newcommand{\lgray}[1]{\color{lightgray}{#1}}\) \(\newcommand{\rank}{\operatorname{rank}}\) \(\newcommand{\row}{\text{Row}}\) \(\newcommand{\col}{\text{Col}}\) \(\renewcommand{\row}{\text{Row}}\) \(\newcommand{\nul}{\text{Nul}}\) \(\newcommand{\var}{\text{Var}}\) \(\newcommand{\corr}{\text{corr}}\) \(\newcommand{\len}[1]{\left|#1\right|}\) \(\newcommand{\bbar}{\overline{\bvec}}\) \(\newcommand{\bhat}{\widehat{\bvec}}\) \(\newcommand{\bperp}{\bvec^\perp}\) \(\newcommand{\xhat}{\widehat{\xvec}}\) \(\newcommand{\vhat}{\widehat{\vvec}}\) \(\newcommand{\uhat}{\widehat{\uvec}}\) \(\newcommand{\what}{\widehat{\wvec}}\) \(\newcommand{\Sighat}{\widehat{\Sigma}}\) \(\newcommand{\lt}{<}\) \(\newcommand{\gt}{>}\) \(\newcommand{\amp}{&}\) \(\definecolor{fillinmathshade}{gray}{0.9}\)Given that a researcher has set a null and alternative hypothesis, has observed data from a population based on a sample, and has measured the amount of evidence against the null hypothesis using a test statistic, it is now time for the researcher make a conclusion. The main question that the researcher faces in this situation is whether enough evidence exists as measured by the test statistic for the researcher to reject the null hypothesis.
In the case of the sample of five exams, suppose as we did before that the mean of the sample is 84.4. This is consistent with the form of the alternative hypothesis, but as we discussed in the previous section, we also know that the mean computed on the sample is an estimate of the population mean, and there will be some error associated with the estimate. If the professor does not know the population mean (and in practical cases it will be unknown), there are two possible decisions the professor can make, and therefore four possible outcomes depending on whether the correct decision has been made (see Table 12.1).
Table 12.1 The four possible outcomes for a statistical hypothesis test.
|
Decision |
\(H_0\) True |
\(H_0\) False |
|
Reject \(H_0\) |
Type I Error |
Correct Decision |
|
Fail to Reject \(H_0\) |
Correct Decision |
Type II Error |
The first two possibilities can occur if the professor decides there is sufficient evidence to reject the null hypothesis. That is, the professor concludes that the mean grade on all the exams is greater than 80. In drawing this conclusion, it is possible that the professor is making the right choice, that the mean really is greater than 80. However, because the professor is deciding based on a sample, it is also possible that the professor is drawing the wrong conclusion, that is, that the mean is not greater than 80. Because of the sample that happened to be selected, the conclusion was made that it is not greater than 80. In terms of statistical hypothesis testing, this is equivalent to rejecting the null hypothesis when it is true. This is known as a Type I Error.
A Type I Error occurs when the null hypothesis is rejected when it is true.
The next two possibilities occur when the professor concludes that the mean grade on all the exams is less than or equal to 80. In drawing this conclusion, it is possible that the professor is making the right choice, that the mean really is less than or equal to 80. However, it is also possible that the professor is drawing the wrong conclusion, that is that the mean is not less than or equal to 80. In terms of statistical hypothesis testing, this is equivalent to failing to reject the null hypothesis when it is false. This is known as a Type II Error.
A Type II Error occurs when the null hypothesis is not rejected when it is false.
Of course, the professor, like any researcher, would like to avoid making either a Type I or Type II Error. But, unless the researcher is willing to do a census of the entire population, this is not possible, and so researchers have to come to terms with the fact that these errors are always possible. However, it should be noted that once a conclusion is made, the type of potential error will be known. If a researcher rejects the null hypothesis, they know that they are either making a correct decision or they are committing a Type I Error because a Type II Error can only occur when the researcher fails to reject the null hypothesis. Similarly, if the researcher fails to reject the null hypothesis, they know that they are either making the correct decision or that they are committing a Type II Error because a Type I Error only occurs when the null hypothesis is rejected.
Knowing that the potential for these errors cannot be avoided completely, the next best option is to attempt to ensure that they do not happen very often. As it turns out, researchers can control how often one of the errors occurs but not both, unless the researcher has a great amount of flexibility with the size of the sample. Let us assume for the moment that a sample size has been decided on, and in that case, we can control how often one of the errors occurs. Given the choice, which error do we want to control? The clear winner for this choice is the Type I Error.
When structuring the null and alternative hypothesis, researchers will say the hypothesis they are trying to prove is the alternative hypothesis. Therefore, if they would like to prove their theory correct, they are interested in rejecting the null hypothesis, corresponding to the first row of Table 12.1. The only error that can occur when rejecting the null hypothesis is a Type I Error. Hence, the researcher has a keen interest in controlling how often this error occurs.
What the researcher will do is specify how much risk they are willing to take with a Type I Error occurring when they analyze their data. This risk is mitigated by specifying the probability that a Type I Error will occur. This value is called the significance level of the test, and it is usually denoted using the Greek letter alpha (\(\alpha\)).
The significance level of a statistical hypothesis test is the probability that the null hypothesis will be rejected when the null hypothesis is true.
Once the significance level is set, if the researcher rejects the null hypothesis, then they are able to say that they found sufficient evidence not to believe the null hypothesis, but then they are also able to acknowledge that they are taking an \(\alpha\) amount of risk in drawing this conclusion.
Usually, the significance levels are set to \(\alpha=0.05\) or \(\alpha=0.01\). The level of \(\alpha=0.05\), also called the 5% significance level, is considered the standard level, but this value is largely arbitrary. This value was suggested by Sir Ronald S. Fisher in the 1930 book Statistical Methods for Research Workers (Cowles and Davis 1982; Fisher 1930).
The question now is why not just set the significance level to be very small, say \(\alpha=0.00000001\)? That way, a researcher will be very sure that they will probably never commit a Type I Error. The problem with this approach is that to ensure the Type I Error rate is very small, the testing procedure will just end up rejecting the null hypothesis only when there is so much evidence against it that there is no other choice. That is, the researcher will rarely, if ever, reject the null hypothesis. When the null hypothesis is rejected, it will only be in the most obvious cases. The problem with this approach is that it goes well beyond the concept of reasonable doubt. Under this type of procedure many potential important discoveries will be ignored because of the researcher’s insistence on never committing a Type I Error.
The other effect of choosing a very small significance level is that because the testing procedure will reject the null hypothesis less often, the test will usually fail to reject the null hypothesis—in this case, the rate at which the test fails to reject the null hypothesis when it should have will become larger. That is, the corresponding Type II Error rate will increase.
Hence, the choice of the significance level results from the need to balance the size of the probabilities of the Type I and Type II Errors. In the way that a statistical test is set up, if a researcher specifies the Type I Error rate, they must accept whatever Type II Error rate is associated with it. What makes this more difficult is the fact that the Type II Error will be unknown. Because of these factors statisticians have traditionally used \(\alpha=0.05\) or \(\alpha=0.01\) as a compromise.
The ability for a statistical testing procedure to detect a true alternative hypothesis is called the power of the test. Specifically, power refers to the probability that a true alternative hypothesis is detected by the test.
The power of a statistical hypothesis test is the probability of rejecting the null hypothesis when the alternative hypothesis is true.
Note that when the null hypothesis is false, there are two possible outcomes: the correct decision (whose probability is equal to the power of the statistical test), or the Type II Error. If the alternative hypothesis is true, one of these outcomes must occur and it follows that the probability of rejecting the null hypothesis when it is false is equal to one minus the probability of a Type II Error.
Therefore a high power corresponds to a low Type II Error. As we stated above, the probability of a Type II Error is set when the significance level is set and is unknown. The only way to ensure that the power of a test will be large, and the Type II Error probability is small, is to choose a large sample from the population. Some researchers will do a power analysis. This is a statistical calculation that tells the researcher about how large the sample needs to be for the test to have certain power characteristics. These calculations can be quite complicated, but a power analysis is usually a sign of a carefully designed study.
Once the significance level is chosen, the test procedure will identify how much evidence is required to reject the null hypothesis. Specifically, the choice of the significance level will identify a range of values for the test statistic for which the researcher would reject the null hypothesis. This range is known as the rejection region.
The rejection region of a statistical test is a range of values such that if the test statistic is in the range, the null hypothesis is rejected for the specified significance level.
For the example of the five exams sampled from the class papers, the test statistic was be written as
\[\frac{\text{sample mean}−80}{\text{standard error}} \nonumber \]
or the number of standard errors the mean of the sample was above the mean specified in the null hypothesis, which was 80. Suppose that a researcher specified the significance level \(\alpha=0.05\) for this test. Some statistical calculations would show that for this significance level the null hypothesis should be rejected when the test statistic is greater that 2.32. This would be the rejection region for the test.
When a null hypothesis is rejected, the researchers will state that the result is significant. So, in the above example, if the test statistic was in the rejection region, the professor would say that the result is significant, meaning that the null hypothesis was rejected. It is also common for researchers to state that “there is a significant difference” when rejecting a null hypothesis. This is used when the null hypothesis refers to the equality between two groups. For example, in a study of income and race or ethnicity, a researcher might state “there was a significant difference between the median income of African American families and white families.”
When talking about statistical results, it is crucial to only say that something is significant when the conclusion is based on rejecting the null hypothesis of a statistical test. Statistical researchers always automatically assume that the term significant relates to a statistical test.

