Skip to main content
Statistics LibreTexts

5.4: Binomial Distribution

  • Page ID
    10936
  • \( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)

    \( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)

    \( \newcommand{\dsum}{\displaystyle\sum\limits} \)

    \( \newcommand{\dint}{\displaystyle\int\limits} \)

    \( \newcommand{\dlim}{\displaystyle\lim\limits} \)

    \( \newcommand{\id}{\mathrm{id}}\) \( \newcommand{\Span}{\mathrm{span}}\)

    ( \newcommand{\kernel}{\mathrm{null}\,}\) \( \newcommand{\range}{\mathrm{range}\,}\)

    \( \newcommand{\RealPart}{\mathrm{Re}}\) \( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)

    \( \newcommand{\Argument}{\mathrm{Arg}}\) \( \newcommand{\norm}[1]{\| #1 \|}\)

    \( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)

    \( \newcommand{\Span}{\mathrm{span}}\)

    \( \newcommand{\id}{\mathrm{id}}\)

    \( \newcommand{\Span}{\mathrm{span}}\)

    \( \newcommand{\kernel}{\mathrm{null}\,}\)

    \( \newcommand{\range}{\mathrm{range}\,}\)

    \( \newcommand{\RealPart}{\mathrm{Re}}\)

    \( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)

    \( \newcommand{\Argument}{\mathrm{Arg}}\)

    \( \newcommand{\norm}[1]{\| #1 \|}\)

    \( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)

    \( \newcommand{\Span}{\mathrm{span}}\) \( \newcommand{\AA}{\unicode[.8,0]{x212B}}\)

    \( \newcommand{\vectorA}[1]{\vec{#1}}      % arrow\)

    \( \newcommand{\vectorAt}[1]{\vec{\text{#1}}}      % arrow\)

    \( \newcommand{\vectorB}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)

    \( \newcommand{\vectorC}[1]{\textbf{#1}} \)

    \( \newcommand{\vectorD}[1]{\overrightarrow{#1}} \)

    \( \newcommand{\vectorDt}[1]{\overrightarrow{\text{#1}}} \)

    \( \newcommand{\vectE}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash{\mathbf {#1}}}} \)

    \( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)

    \(\newcommand{\longvect}{\overrightarrow}\)

    \( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)

    \(\newcommand{\avec}{\mathbf a}\) \(\newcommand{\bvec}{\mathbf b}\) \(\newcommand{\cvec}{\mathbf c}\) \(\newcommand{\dvec}{\mathbf d}\) \(\newcommand{\dtil}{\widetilde{\mathbf d}}\) \(\newcommand{\evec}{\mathbf e}\) \(\newcommand{\fvec}{\mathbf f}\) \(\newcommand{\nvec}{\mathbf n}\) \(\newcommand{\pvec}{\mathbf p}\) \(\newcommand{\qvec}{\mathbf q}\) \(\newcommand{\svec}{\mathbf s}\) \(\newcommand{\tvec}{\mathbf t}\) \(\newcommand{\uvec}{\mathbf u}\) \(\newcommand{\vvec}{\mathbf v}\) \(\newcommand{\wvec}{\mathbf w}\) \(\newcommand{\xvec}{\mathbf x}\) \(\newcommand{\yvec}{\mathbf y}\) \(\newcommand{\zvec}{\mathbf z}\) \(\newcommand{\rvec}{\mathbf r}\) \(\newcommand{\mvec}{\mathbf m}\) \(\newcommand{\zerovec}{\mathbf 0}\) \(\newcommand{\onevec}{\mathbf 1}\) \(\newcommand{\real}{\mathbb R}\) \(\newcommand{\twovec}[2]{\left[\begin{array}{r}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\ctwovec}[2]{\left[\begin{array}{c}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\threevec}[3]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\cthreevec}[3]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\fourvec}[4]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\cfourvec}[4]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\fivevec}[5]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\cfivevec}[5]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\mattwo}[4]{\left[\begin{array}{rr}#1 \amp #2 \\ #3 \amp #4 \\ \end{array}\right]}\) \(\newcommand{\laspan}[1]{\text{Span}\{#1\}}\) \(\newcommand{\bcal}{\cal B}\) \(\newcommand{\ccal}{\cal C}\) \(\newcommand{\scal}{\cal S}\) \(\newcommand{\wcal}{\cal W}\) \(\newcommand{\ecal}{\cal E}\) \(\newcommand{\coords}[2]{\left\{#1\right\}_{#2}}\) \(\newcommand{\gray}[1]{\color{gray}{#1}}\) \(\newcommand{\lgray}[1]{\color{lightgray}{#1}}\) \(\newcommand{\rank}{\operatorname{rank}}\) \(\newcommand{\row}{\text{Row}}\) \(\newcommand{\col}{\text{Col}}\) \(\renewcommand{\row}{\text{Row}}\) \(\newcommand{\nul}{\text{Nul}}\) \(\newcommand{\var}{\text{Var}}\) \(\newcommand{\corr}{\text{corr}}\) \(\newcommand{\len}[1]{\left|#1\right|}\) \(\newcommand{\bbar}{\overline{\bvec}}\) \(\newcommand{\bhat}{\widehat{\bvec}}\) \(\newcommand{\bperp}{\bvec^\perp}\) \(\newcommand{\xhat}{\widehat{\xvec}}\) \(\newcommand{\vhat}{\widehat{\vvec}}\) \(\newcommand{\uhat}{\widehat{\uvec}}\) \(\newcommand{\what}{\widehat{\wvec}}\) \(\newcommand{\Sighat}{\widehat{\Sigma}}\) \(\newcommand{\lt}{<}\) \(\newcommand{\gt}{>}\) \(\newcommand{\amp}{&}\) \(\definecolor{fillinmathshade}{gray}{0.9}\)

    Introduction

    Recall that when calculating the probability of multiple discrete events occurring we used the formula

    \[\text{P(A and B)} =  \text{P(A)} \cdot \text{P(B|A)} \label{eq:5.4.1.1}\]

    If we know that A and B are independent events, then \(\text{P(B)} = \text{P(B|A)}\).  So, for more than two independent events to occur, we are simply multiplying their probabilities together as demonstrated in the next example. 

    Example \(\PageIndex{1}\): Shock Study

    Suppose we randomly selected four individuals to participate in the "shock" study. Let's call the four people Allen (A), Brittany (B), Caroline (C), and Damian (D) for convenience. They have the option to refuse to participate or to give a severe shock. Only 35% of people refuse. What is the chance that exactly one of the people refuses? 

    Solution

    There are four events and each event only has two outcomes: refuse or shock. Since one person's choice is not going to affect another person's choice, these events are independent.   

    We are interested in refusals, so let's call that our "success" case. The other outcome, shock, is the complement: \(1 - 0.35 = 0.65\). This will be a "failure" case for our question.

    First, let's consider a scenario where one person refuses:

    \[ \begin{align*} P(A = \text{refuse}; B = \text{shock}; C = \text{shock}; D = \text{shock}) &= P(A = \text{refuse}) P(B = \text{shock}) P(C = \text{shock}) P(D = \text{shock}) \\[5pt] &= (0.35)(0.65)(0.65)(0.65) \\[5pt] &= (0.35)^1(0.65)^3 \\[5pt] &= 0.096 \end{align*}\]

    But there are three other scenarios: Brittany, Caroline, or Damian could have been the one to refuse. In each of these cases, the probability is again

    \[P=(0.35)^1(0.65)^3. \nonumber\]

    These four scenarios exhaust all the possible ways that exactly one of these four people could refuse to administer the most severe shock, so the total probability is

    \[P(\text{Exactly One Person Refuses}) = 4 \times (0.35)^1(0.65)^3 = 0.38. \nonumber\]

    Exercise \(\PageIndex{1}\)

    Verify that the scenario where Brittany is the only one to refuse to give the most severe shock has probability \((0.35)^1(0.65)^3.\)

    Answer

    \[ \begin{align*} P(A = shock; B = refuse; C = shock; D = shock) &= (0.65)(0.35)(0.65)(0.65) \\[5pt] &= (0.35)^1(0.65)^3.\end{align*}\]

    The Binomial Distribution

    The scenario outlined in Example \(\PageIndex{1}\) is a special case the Multiplication rule that uses the binomial distribution. A binomial distribution describes the probability of having exactly \(x = r\) successes in \(n\) independent Bernoulli trials with probability of a success \(p\) (in Example \(\PageIndex{1}\), \(n = 4, r = 1, p = 0.35\) and \(x=\text{the number of people who refuse}\) as our discrete random variable).

    We would like to determine the probabilities associated with the binomial distribution more generally, i.e. we want a formula where we can use \(n\), \(x=r\), and \(p\) to obtain the probability. To do this, we reexamine each part of the example.

    There were four individuals who could have been the one to refuse, and each of these four scenarios had the same probability. Thus, we could identify the final probability as

    \[ \text {(Number of scenarios)} \times \text {P(single scenario)} \label{eq:5.4.1.2}\]

    The first component of this equation is the number of ways to arrange the \(r = 1\) successes among the \(n = 4\) trials. The second component is the probability of any of the four (equally probable) scenarios.

    Consider P(single scenario) under the general case of \(r\) successes and \(n - r\) failures in the \(n\) trials. In any such scenario, we apply the Multiplication Rule for independent events:

    \[P(x=r) = p^{r} (1- p)^{n-r} \label{eq:5.4.1.3}\]

    This is our general formula for P(single scenario) where there are only two outcomes.

    Secondly, we introduce a general formula for the number of ways to choose \(r\) successes in \(n\) trials, i.e. arrange \(r\) successes and \(n - r\) failures:

    \[ \binom {n}{r} = \dfrac {n!}{r!(n - r)!} \label{eq:5.4.1.4} \]

    The quantity \( \binom {n}{r} \) is read n choose r. Remember that the exclamation point notation (e.g. r!) denotes a factorial expression.

    \[ \begin{align} 0! &= 1 \nonumber \\[5pt] 1! &= 1 \nonumber \\[5pt] 2! &= 2 \times 1 = 2 \nonumber \\[5pt] 3! &= 3 \times 2 \times 1 = 6 \nonumber \\[5pt] 4! &= 4 \times 3 \times 2 \times 1 = 24 \nonumber \\[5pt] & \vdots \nonumber \\[5pt] n! &= n \times (n - 1) \times \dots \times 3 \times 2 \times 1 \label{eq:5.4.1.5} \end{align} \]

    Substituting Equation \ref{eq:5.4.1.5} into Equation \ref{eq:5.4.1.4}, we can compute the number of ways to choose \(r = 1\) successes in \(n = 4\) trials:

    \[ \begin{align} \binom {4}{1} &= \dfrac {4!}{1! (4 - 1)!} \nonumber\\[5pt]  &= \dfrac {4!}{1! 3!} \nonumber\\[5pt]&= \dfrac {4 \times 3 \times 2 \times 1}{(1)(3 \times 2 \times 1)} \nonumber\\[5pt]&= 4 \nonumber\end{align} \]

    This result is exactly what we found by carefully thinking of each possible scenario in Example \(\PageIndex{1}\).

    Other notations

    Other notation for n choose r includes \(nC_r\), \(C^r_n\), and \(C(n, r)\).

    Substituting \(n\) choose \(r\) for the number of scenarios and \(p^{r}(1 - p)^{n-r}\) for the single scenario probability in Equation \ref{eq:5.4.1.2} yields the general binomial formula (Equation \ref{eq:5.4.1.6}).

    Definition: Binomial distribution

    Suppose the probability of a single trial being a success is p. Then the probability of observing exactly r successes in n independent trials is given by

    \[P(x = r) = \binom {n}{r} p^r (1 - p)^{n - r} = \dfrac {n!}{r!(n - r)!} p^r (1 - p)^{n - r} \label{eq:5.4.1.6} \]

    TIP: Four conditions must be met to be a binomial distribution

    The probability for discrete random variables can be calculated using a binomial distribution when

    1. The number of trials, \(n\), is fixed.
    2. Each trial outcome can be classified as a success or failure.
    3. The trials independent.
    4. The probability of a success, \(p\), is the same for each trial.

    Note: Conditions 3 and 4 could be considered one condition since independence implies that the probability from one event to the next will remain the same.

    Example \(\PageIndex{2}\)

    Using the same scenario from Example \(\PageIndex{1}\), but this time 8 people are randomly selected. What is the probability that 3 of 8 randomly selected students will refuse to administer the worst shock?

    Solution

    We would like to apply the binomial model, so we check our conditions.

    Since we can only apply the binomial model to discrete variables, start there. We are counting the number of people who will refuse. Since we can't have partial people, the variable is discrete: x = the number of people who refuse. Now, on to the conditions.

    Condition 1: The number of trials is fixed (n = 8)

    Condition 2: Each trial outcome can be classified as a success or failure

    Condition 3: Because the sample is random, the trials are independent

    Condition 4: The probability of a success is the same for each trial since the trials are independent. 

    In the outcome of interest, there are \(r = 3\) successes in \(n = 8\) trials, and the probability of a success is \(p = 0.35\).

    So the probability that 3 of 8 will refuse is given by

    \[\begin{align*}P(x = 3) =  \binom {8}{3} {(0.35)}^3 (1 - 0.35)^{8 - 3} &= \dfrac {8!}{3!(8 - 3)!} {(0.35)}^3 (1 - 0.35)^{8 - 3} \\[5pt] &= \dfrac {8!}{3! 5!} {(0.35)}^3 {(0.65)}^5 \end{align*}\]

    Dealing with the factorial part:

    \[ \begin{align*} \dfrac {8!}{3!5!} &= \dfrac {8 \times 7 \times 6 \times 5 \times 4 \times 3 \times 2 \times 1}{(3 \times 2 \times 1)( 5 \times 4 \times 3 \times 2 \times 1 )} \\[5pt] &= \dfrac {8 \times 7 \times 6 }{ 3 \times 2 \times 1 } \\[5pt] & = 56 \end{align*}\]

    Using \((0.35)^3(0.65)^5 \approx 0.005\), the final probability is about \(P(x = 3) = 56 \cdot 0.005 = 0.28\).

    TIP: Computing binomial probabilities

    The first step in using the binomial model is to check that the model is appropriate.

    The second step is to identify n, p, and r.

    The final step is to apply the formulas and interpret the results.

    Independence

    As a reminder, the condition for independence could be satisfied in a few ways. 

    1. The events do not affect each other. 

    Examples of this include guessing on answers for a multiple choice test, rolling a fair die, flipping a coin, responses from unassociated people, etc. 

    1. Possible outcomes are replaced after a selection is made. 

    Examples of this include picking an object out of a container and putting it back before the next selection, allowing people to be selected only once, etc. 

    1. The sample size is 5% or less of the total population, then selections can be treated as independent events since the removal of an outcome doesn't overly affect the final calculation. 

    An example of this is randomly sampling 200 people and asking the question, what is the probability that 8 of them completed high school. 5% of 200 is 10. Since 8 is less than 10, we could treat our calculations as independent even though we are not replacing the individuals.  

    Another example is being told that 51% of Americans have a valid U.S. passport and being asked to determine the probability of selecting 100 Americans with a valid passport. The population in the U.S. is 342.6 million as of 2026. 100 is far less than 5% of the population so we can treat each selection as independent to make calculations easier.

    Example \(\PageIndex{3}\)

    The probability that a random smoker will develop a severe lung condition in his or her lifetime is about 0.3. If you have 4 friends who smoke, are the conditions for the binomial model satisfied?

    Solution

    One possible answer: If the friends know each other, then the independence assumption is probably not satisfied. For example, acquaintances may have similar smoking habits.

    Exercise \(\PageIndex{2}\)

    A lacrosse team is selecting a captain. The names of all the seniors are put into a hat, and the first three that are drawn will be the captains. The names are not replaced once they are drawn (one person cannot be two captains). You want to see if the captains all play the same position. State whether this is binomial or not and state why.

    Answer

    This is not binomial because the names are not replaced, which means the probability changes for each time a name is drawn. This violates the condition of independence.

     Calculating Binomial Probabilities of more than one event

    Using the Shock Study example, where 35% of participants refuse to administer a shock and 8 people are randomly selected for the study, let's create a probability distribution table.  We know that \(n=8\), \(p=0.35\), and \(x =\) 0, 1, 2, 3 ,4 ,5 ,6 ,7 or 8 since we could any anywhere from 0 successes to 8 successes. For our distribution, we will use our binomial formula for each value of x. 

    Shock Study - Binomial Probabilities
    x = # of people who refuse P(X)
    0 \(P(x = 0) =  \binom {8}{0} {(0.35)}^0 {(0.65)}^8 = 0.0318...\approx 0.03\)
    1 \(P(x = 1) =  \binom {8}{1} {(0.35)}^1 {(0.65)}^7 = 0.1372...\approx 0.14\)
    2 \(P(x = 2) =  \binom {8}{2} {(0.35)}^2 {(0.65)}^6 =0.2586... \approx 0.26\)
    3 \(P(x = 3) =  \binom {8}{3} {(0.35)}^3 {(0.65)}^5 = 0.2785...\approx 0.28\)
    4 \(P(x = 4) =  \binom {8}{4} {(0.35)}^4 {(0.65)}^4 = 0.1875...\approx 0.19\)
    5 \(P(x = 5) =  \binom {8}{5} {(0.35)}^5 {(0.65)}^3 = 0.0807... \approx 0.08\)
    6 \(P(x = 6) =  \binom {8}{6} {(0.35)}^6{(0.65)}^5 = 0.0217...\approx 0.02\)
    7 \(P(x = 7) =  \binom {8}{7} {(0.35)}^7 {(0.65)}^1 = 0.0033...\approx 0.003\)
    8 \(P(x = 8) =  \binom {8}{8} {(0.35)}^8 {(0.65)}^0 = 0.0002...\approx 0.0002\)

    Notice that while we could round the probabilities for \(x = 0\) through \(x = 6\) to two decimal places, \(x = 7\) and \(x = 8\) would round to 0, implying that those events are impossible. However, just a few more decimal places out we see non-zero values. It is good practice to round to the first non-zero number in these cases. 

    Additionally, we can expect that the probability values in this distribution will sum to 1 since it accounts for all possible outcomes for this study. 

    That is 

    \[P(x = 0) + P(x = 1)+ P(x = 2)+ P(x = 3)+ P(x = 4)+ P(x = 5)+ P(x = 6)+ P(x = 7)+ P(x = 8) = 1 \nonumber\]

    If we want to calculate the probability that more than 5 people will refuse, then what we are asking is to find \(P(x > 5)\). Just like with earlier distributions, we can add the probability values for the classes that meet this condition. To avoid a round off error, use the unrounded values in your calculation.

    \[P(x > 5) = P(x = 6) + P(x = 7) + P(x = 8) = 0.0217...+0.0033...0.0002...= 0.02532... \approx = 0.025 \nonumber\]

    Since this type of distribution captures all of the possible outcomes, we can also use the complement in calculations. For example, if we want to find the probability that 5 or fewer people will refuse, then we are asking for \(P(x \leq 5)\). 

    \[P(x \leq 5) + P(x > 5) = 1 \nonumber \]

    \[P(x \leq 5) = 1 - P(x > 5) \nonumber\]

    \[P(x \leq 5) = 1 - 0.02532... \nonumber\]

    \[P(x \leq 5) = 0.9746... \nonumber\]

    \[P(x \leq 5) \approx 0.975 \nonumber\]

    Example \(\PageIndex{4}\)

    Suppose these four friends do not know each other and we can treat them as if they were a random sample from the population. The probability that a random smoker will develop a severe lung condition in his or her lifetime is about 0.3. Is the binomial model appropriate? What is the probability that

    1. none of them will develop a severe lung condition?
    2. One will develop a severe lung condition?
    3. That no more than one will develop a severe lung condition?

    Solution

    We are counting people, so we have a discrete random variable. Let's define it as x = the number of people who develop a lung condition. 

    To check if the binomial model is appropriate, we must verify the conditions.

    Condition 1: We have a fixed number of trials (n = 4).

    Condition 2: Each outcome is a success or failure.

    Condition 3: Since we are supposing we can treat the friends as a random sample, they are independent.

    Condition 4: The probability of a success is the same for each trials since the individuals are like a random sample (p = 0.3 if we say a "success" is someone getting a lung condition, a morbid choice).

    Compute parts (a) and (b) from the binomial formula in Equation \ref{eq:5.4.1.6}:

    Part (a) \[P(x = 0) = \binom {4}{0}(0.3)^0(0.7)^4 = 1 \times 1 \times 0.7^4 = 0.2401 \nonumber\]

    Note: 0! = 1.

    There is about a 24% chance that none of them will develop a severe lung condition.

    Part (b) \[P(x = 1) = \binom {4}{1}(0.3)^1(0.7)^3 = 0.4116. \nonumber\]

    There is about a 41% chance that one of them will develop a severe lung condition.

    Part (c) can be computed as the sum of parts (a) and (b):

    \[P(x \leq 1) = P(x = 0)+P(x = 1) = 0.2401+0.4116 = 0.6517. \nonumber\]

    That is, there is about a 65% chance that no more than one of your four smoking friends will develop a severe lung condition.

    Exercise \(\PageIndex{3}\)

    The probability that a random smoker will develop a severe lung condition in his or her lifetime is about 0.3. Still assuming we can treat this sample of 4 as independent samples, what is the probability that at least 2 of your 4 smoking friends will develop a severe lung condition in their lifetimes?

    Answer

    Since n = 4, the binomial distribution has possibilities of x = 0, 1, 2, 3 or 4. In Example \(\PageIndex{4}\), we computer \(P(x \leq 1)\) to be 0.6517. "At least 2" is 2, 3 or 4, the rest of the possibilities, so we can use the complement to calculate instead of adding the probability values for 2, 3, and 4.

    \[P(x \geq 2) = 1 - P(x \leq 1)\nonumber\]

    \[P(x \geq 2) = 1 - 0.6517\nonumber\]

    \[P(x \geq 2) = 0.3483 \nonumber\]

    There is about a 35% chance that at least 2 of the 4 smoking friends will develop a severe lung condition in their lifetimes. 

    TIP: computing n choose r

    In general, it is useful to do some cancelation in the factorials immediately if calculating mostly by hand.

    Alternatively, many computer programs and calculators have built in functions to compute n choose r, factorials, and even entire binomial probabilities. 

    Check with your instructor for guidance on technology use.

    Measures of Center and Variation for Binomial Distribution

    Recall that for probability distribution, the mean, μ, of a discrete probability function is the expected value. 

    \[μ=∑(x \cdot P(x))\nonumber\]

    The standard deviation, σ, is the square root of the variance.

    \[σ=\sqrt{∑[(x – μ)2 \cdot P(x)]}\nonumber\]

    However, if we know that the distribution is binomial, then the mean, variance, and standard deviation of the number of observed successes can be calculated with more ease:

    Mean: \[\mu = np \label{eq:5.4.1.7}\] 

    Variance: \[\sigma^2 = np(1 - p) \label{eq:5.4.1.8}\]

    Standard Deviation: \[\sigma = \sqrt {np(1- p)} \label{eq:5.4.1.9} \]

     

    Example \(\PageIndex{5}\)

    If you ran a study and randomly sampled 40 students, how many would you expect to refuse to administer the worst shock? What is the standard deviation of the number of people who would refuse? 

    Solution

    We are asked to determine the expected number (the mean) and the standard deviation.

    \[\mu = np = 40 \times 0.35 = 14 \nonumber\]

    and

    \[\sigma = \sqrt{np(1 - p)} = \sqrt { 40 \times 0.35 \times 0.65} = 0.02. \nonumber\]

    Because very roughly 95% of observations fall within 2 standard deviations of the mean, we would probably observe at least 8 but less than 20 individuals in our sample who would refuse to administer the shock.

     

    Exercise \(\PageIndex{4}\)

    Suppose you have 7 friends who are smokers and they can be treated as a random sample of smokers. As with earlier examples, the probability that a random smoker will develop a severe lung condition in his or her lifetime is about 0.3.

    1. How many would you expect to develop a severe lung condition, i.e. what is the mean?
    2. What is the probability that at most 2 of your 7 friends will develop a severe lung condition.
    Answer a

    \(\mu\) = 0.3(7) = 2.1.

    We expect 2 of the 7 friends to develop a severe lung condition. 

    Answer b

    Let's define our discrete random variable as x = the number of people who develop a severe lung condition. 

    "At most" means 0, 1, or 2 people, in other words \(x \leq 2\).

    \(P(x \leq 2) = P(x = 0)+P(x = 1)+P(x = 2) = 0.6471.\)

    The probability that at most 2 of the 7 friends will develop a severe lung condition is about 65%.

    Something interesting about terms in the binomial distribution

    Below we consider the first term in the binomial probability, n choose k under some special scenarios.

    Example \(\PageIndex{6}\)

    Why is it true that \( \binom {n}{0} = 1\) and \( \binom {n}{n} = 1 \) for any number n?

    Solution

    Frame these expressions into words.

    How many different ways are there to arrange 0 successes and n failures in n trials? (1 way.)

    How many different ways are there to arrange n successes and 0 failures in n trials? (1 way.)

    If we do the math, the result is the same, but intuition in these cases works just as well. 

    \[\binom {n}{0} = \dfrac{n!}{(n-0)!\cdot 0!} = \dfrac{n!}{n!\cdot 0!}= 1 \nonumber\]

    and

    \[\binom {n}{n} = \dfrac{n!}{(n-n)!\cdot n!} = \dfrac{n!}{0!\cdot n!} = 1 \nonumber\]

    Note: 0! = 1 and \(\dfrac{n!}{n!}=1\) since a number divided by itself is 1.

    Example \(\PageIndex{5}\)

    How many ways can you arrange one success and n -1 failures in n trials? How many ways can you arrange n -1 successes and one failure in n trials?

    Solution

    One success and n - 1 failures: there are exactly n unique places we can put the success, so there are n ways to arrange one success and n - 1 failures.

    A similar argument is used for the second question.

    Mathematically, we show these results by verifying the following two equations:

    \[ \binom {n}{1} = \dfrac{n!}{(n-1)!\cdot 1!}=\dfrac{n \cdot (n-1) \cdot ... \cdot 1}{(n-1) \cdot ... \cdot 1} = n \nonumber\]

    and

    \[ \binom {n}{n - 1} = \dfrac{n!}{(n-(n-1))!\cdot (n-1)!}= \dfrac{n!}{1!\cdot (n-1)!} = \dfrac{n \cdot (n-1) \cdot ... \cdot 1}{(n-1) \cdot ... \cdot 1} = n \nonumber\]

    Technology and the Binomial Distribution

    Depending on your course and instructor, you may be allowed to use technology in the form of binomial distribution functions on your calculator or a spreadsheet type program. Before using these functions, it is important to have a good grasp on the concept of the probabilities that are involved in finding the requested probability. So, please check with your instructor be taking these shortcuts and make sure you are comfortable with the concept before relying too heavily on the technology.

    Notation for the Binomial: \(B =\) Binomial Probability Distribution Function

    \[X \sim B(n, p)\]

    Read this as "\(X\) is a random variable with a binomial distribution." The parameters are \(n\) and \(p\); \(n =\) number of trials, \(p =\) probability of a success on each trial.

    Example \(\PageIndex{8}\)

    It has been stated that about 41% of adult workers have a high school diploma but do not pursue any further education. If 20 adult workers are randomly selected, find the probability that at most 12 of them have a high school diploma but do not pursue any further education. How many adult workers do you expect to have a high school diploma but do not pursue any further education?

    Let \(X\) = the number of workers who have a high school diploma but do not pursue any further education.

    \(X\) takes on the values 0, 1, 2, ..., 20 where \(n = 20, p = 0.41\), and \(q = 1 – 0.41 = 0.59\). \(X \sim B(20, 0.41)\)

    Find \(P(x \leq 12)\). \(P(x \leq 12) = 0.9738\). (calculator (TI83/84) or computer)

    Go into 2nd DISTR. The syntax for the instructions are as follows:

    To calculate (\(x = \text{value}): \text{binompdf}(n, p, \text{number}\)) if "number" is left out, the result is the binomial probability table.

    To calculate \(P(x \leq \text{value}): \text{binomcdf}(n, p, \text{number})\) if "number" is left out, the result is the cumulative binomial probability table.

    For this problem: After you are in 2nd DISTR, arrow down to binomcdf. Press ENTER. Enter 20,0.41,12). The result is \(P(x \leq 12) = 0.9738\).

    If you want to find \(P(x = 12)\), use the pdf (binompdf). If you want to find \(P(x > 12)\), use \(1 - \text{binomcdf}(20,0.41,12)\).

    The probability that at most 12 workers have a high school diploma but do not pursue any further education is 0.9738.

    The graph of \(X \sim B(20, 0.41)\) is as follows:

    This histogram shows a binomial probability distribution. It is made up of bars that are fairly normally distributed. The x-axis shows values from 0 to 20. The y-axis shows values from 0 to 0.2 in increments of 0.05.
    Figure \(\PageIndex{1}\) : The graph of \(X \sim B(20, 0.41)\).

    The y-axis contains the probability of \(x\), where \(X =\) the number of workers who have only a high school diploma.

    The number of adult workers that you expect to have a high school diploma but not pursue any further education is the mean, \(\mu = np = (20)(0.41) = 8.2\).

    The formula for the variance is \(\sigma^{2} = npq\). The standard deviation is \(\sigma = \sqrt{npq}\).

    \[\sigma = \sqrt{(20)(0.41)(0.59)} = 2.20.\]

    Exercise \(\PageIndex{6}\)

    About 32% of students participate in a community volunteer program outside of school. If 30 students are selected at random, find the probability that at most 14 of them participate in a community volunteer program outside of school. Use the TI-83+ or TI-84 calculator to find the answer.

    Answer

    \(P(x \leq 14) = 0.9695\)

    Example \(\PageIndex{9}\)

    In the 2013 Jerry’s Artarama art supplies catalog, there are 560 pages. Eight of the pages feature signature artists. Suppose we randomly sample 100 pages. Let \(X =\) the number of pages that feature signature artists.

    1. What values does \(x\) take on?
    2. What is the probability distribution? Find the following probabilities:
      1. the probability that two pages feature signature artists
      2. the probability that at most six pages feature signature artists
      3. the probability that more than three pages feature signature artists.
    3. Using the formulas, calculate the (i) mean and (ii) standard deviation.

    Answer

    1. \(x = 0, 1, 2, 3, 4, 5, 6, 7, 8\)
    2. \(X \sim B(100,8560)(100,8560)\)
      1. \(P(x = 2) = \text{binompdf}\left(100,\dfrac{8}{560},2\right) = 0.2466\)
      2. \(P(x \leq 6) = \text{binomcdf}\left(100,\dfrac{8}{560},6\right) = 0.9994\)
      3. \(P(x > 3) = 1 – P(x \leq 3) = 1 – \text{binomcdf}\left(100,\dfrac{8}{560},3\right) = 1 – 0.9443 = 0.0557\)
      1. Mean \(= np = (100)\left(\dfrac{8}{560}\right) = \dfrac{800}{560} \approx 1.4286\)
      2. Standard Deviation \(= \sqrt{npq} = \sqrt{(100)\left(\dfrac{8}{560}\right)\left(\dfrac{552}{560}\right)} \approx 1.1867\)
    Exercise \(\PageIndex{7}\)

    According to a Gallup poll, 60% of American adults prefer saving over spending. Let \(X\) = the number of American adults out of a random sample of 50 who prefer saving to spending.

    1. What is the probability distribution for \(X\)?
    2. Use your calculator to find the following probabilities:
      1. the probability that 25 adults in the sample prefer saving over spending
      2. the probability that at most 20 adults prefer saving
      3. the probability that more than 30 adults prefer saving
    3. Using the formulas, calculate the (i) mean and (ii) standard deviation of \(X\).
    Answer
    1. \(X \sim B(50, 0.6)\)
    2. Using the TI-83, 83+, 84 calculator with instructions as provided in Example:
      1. \(P(x = 25) = \text{binompdf}(50, 0.6, 25) = 0.0405\)
      2. \(P(x \leq 20) = \text{binomcdf}(50, 0.6, 20) = 0.0034\)
      3. \((x > 30) = 1 - \text{binomcdf}(50, 0.6, 30) = 1 – 0.5535 = 0.4465\)
      1. Mean \(= np = 50(0.6) = 30\)
      2. Standard Deviation \(= \sqrt{npq} = \sqrt{50(0.6)(0.4)} \approx 3.4641\)
    Example \(\PageIndex{10}\)

    The lifetime risk of developing pancreatic cancer is about one in 78 (1.28%). Suppose we randomly sample 200 people. Let \(X\) = the number of people who will develop pancreatic cancer.

    1. What is the probability distribution for \(X\)?
    2. Using the formulas, calculate the (i) mean and (ii) standard deviation of \(X\).
    3. Use your calculator to find the probability that at most eight people develop pancreatic cancer
    4. Is it more likely that five or six people will develop pancreatic cancer? Justify your answer numerically.

    Answer

    1. \(X \sim B(200, 0.0128)\)
      1. Mean \(= np = 200(0.0128) = 2.56\)
      2. Standard Deviation \(= \sqrt{npq} = \sqrt{(200)(0.0128)(0.9872)} \approx 1.5897\)
    2. Using the TI-83, 83+, 84 calculator with instructions as provided in Example:
      \(P(x \leq 8) = \text{binomcdf}(200, 0.0128, 8) = 0.9988\)
    3. \(P(x = 5) = \text{binompdf}(200, 0.0128, 5) = 0.0707\)
      \(P(x = 6) = \text{binompdf}(200, 0.0128, 6) = 0.0298\)
      So \(P(x = 5) > P(x = 6)\); it is more likely that five people will develop cancer than six.
    Exercise \(\PageIndex{8}\)

    During the 2013 regular NBA season, DeAndre Jordan of the Los Angeles Clippers had the highest field goal completion rate in the league. DeAndre scored with 61.3% of his shots. Suppose you choose a random sample of 80 shots made by DeAndre during the 2013 season. Let \(X =\) the number of shots that scored points.

    1. What is the probability distribution for \(X\)?
    2. Using the formulas, calculate the (i) mean and (ii) standard deviation of \(X\).
    3. Use your calculator to find the probability that DeAndre scored with 60 of these shots.
    4. Find the probability that DeAndre scored with more than 50 of these shots.
    Answer
    1. \(X \sim B(80, 0.613)\)
      1. Mean \(= np = 80(0.613) = 49.04\)
      2. Standard Deviation \(= \sqrt{npq} = \sqrt{80(0.613)(0.387)} \approx 4.3564\)
    2. Using the TI-83, 83+, 84 calculator with instructions as provided in Example:
      \(P(x = 60) = \text{binompdf}(80, 0.613, 60) = 0.0036\)
    3. \(P(x > 50) = 1 – P(x \leq 50) = 1 – \text{binomcdf}(80, 0.613, 50) = 1 – 0.6282 = 0.3718\)

    Formula Review

    • \(X \sim B(n, p)\) means that the discrete random variable \(X\) has a binomial probability distribution with \(n\) trials and probability of success \(p\).
    • \(X =\) the number of successes in \(n\) independent trials
    • \(n =\) the number of independent trials
    • \(X\) takes on the values \(x = 0, 1, 2, 3, \dotsc, n\)
    • \(p =\) the probability of a success for any trial
    • \(q =\) the probability of a failure for any trial
    • \(p + q = 1\)
    • \(q = 1 – p\)
    • Suppose the probability of a single trial being a success is p. Then the probability of observing exactly r successes in n independent trials is given by

      \[P(x = r) = \binom {n}{r} p^r (1 - p)^{n - r} = \dfrac {n!}{r!(n - r)!} p^r (1 - p)^{n - r} \nonumber\]

    • The mean of \(X\) is \(\mu = np\). The standard deviation of \(X\) is \(\sigma = \sqrt{npq}\).

    Glossary

    Binomial Experiment
    a statistical experiment that satisfies the following three conditions:
    1. There are a fixed number of trials, \(n\).
    2. There are only two possible outcomes, called "success" and, "failure," for each trial. The letter \(p\) denotes the probability of a success on one trial, and \(q\) denotes the probability of a failure on one trial.
    3. The \(n\) trials are independent and are repeated using identical conditions.
    Bernoulli Trials
    an experiment with the following characteristics:
    1. There are only two possible outcomes called “success” and “failure” for each trial.
    2. The probability \(p\) of a success is the same for any trial (so the probability \(q = 1 − p\) of a failure is the same for any trial).
    Binomial Probability Distribution
    a discrete random variable (RV) that arises from Bernoulli trials; there are a fixed number, \(n\), of independent trials. “Independent” means that the result of any trial (for example, trial one) does not affect the results of the following trials, and all trials are conducted under the same conditions. Under these circumstances the binomial RV \(X\) is defined as the number of successes in \(n\) trials. The notation is: \(X ~ B(n, p)\). The mean is \(\mu = np\) and the standard deviation is \(\sigma = \sqrt{npq}\). The probability of exactly \(x\) successes in \(n\) trials is
    \(P(X = x) = {n \choose x}p^{x}q^{n-x}\).

    Contributors and Attributions

    • David M Diez (Google/YouTube), Christopher D Barr (Harvard School of Public Health), Mine Çetinkaya-Rundel (Duke University)


    This page titled 5.4: Binomial Distribution is shared under a CC BY-SA 3.0 license and was authored, remixed, and/or curated by David Diez, Christopher Barr, and Mine Çetinkaya-Rundel via source content that was edited to the style and standards of the LibreTexts platform.