5.1: Basics of Probability Distributions
- Page ID
- 45478
\( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)
\( \newcommand{\dsum}{\displaystyle\sum\limits} \)
\( \newcommand{\dint}{\displaystyle\int\limits} \)
\( \newcommand{\dlim}{\displaystyle\lim\limits} \)
\( \newcommand{\id}{\mathrm{id}}\) \( \newcommand{\Span}{\mathrm{span}}\)
( \newcommand{\kernel}{\mathrm{null}\,}\) \( \newcommand{\range}{\mathrm{range}\,}\)
\( \newcommand{\RealPart}{\mathrm{Re}}\) \( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)
\( \newcommand{\Argument}{\mathrm{Arg}}\) \( \newcommand{\norm}[1]{\| #1 \|}\)
\( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)
\( \newcommand{\Span}{\mathrm{span}}\)
\( \newcommand{\id}{\mathrm{id}}\)
\( \newcommand{\Span}{\mathrm{span}}\)
\( \newcommand{\kernel}{\mathrm{null}\,}\)
\( \newcommand{\range}{\mathrm{range}\,}\)
\( \newcommand{\RealPart}{\mathrm{Re}}\)
\( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)
\( \newcommand{\Argument}{\mathrm{Arg}}\)
\( \newcommand{\norm}[1]{\| #1 \|}\)
\( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)
\( \newcommand{\Span}{\mathrm{span}}\) \( \newcommand{\AA}{\unicode[.8,0]{x212B}}\)
\( \newcommand{\vectorA}[1]{\vec{#1}} % arrow\)
\( \newcommand{\vectorAt}[1]{\vec{\text{#1}}} % arrow\)
\( \newcommand{\vectorB}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\( \newcommand{\vectorC}[1]{\textbf{#1}} \)
\( \newcommand{\vectorD}[1]{\overrightarrow{#1}} \)
\( \newcommand{\vectorDt}[1]{\overrightarrow{\text{#1}}} \)
\( \newcommand{\vectE}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash{\mathbf {#1}}}} \)
\( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\(\newcommand{\longvect}{\overrightarrow}\)
\( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)
\(\newcommand{\avec}{\mathbf a}\) \(\newcommand{\bvec}{\mathbf b}\) \(\newcommand{\cvec}{\mathbf c}\) \(\newcommand{\dvec}{\mathbf d}\) \(\newcommand{\dtil}{\widetilde{\mathbf d}}\) \(\newcommand{\evec}{\mathbf e}\) \(\newcommand{\fvec}{\mathbf f}\) \(\newcommand{\nvec}{\mathbf n}\) \(\newcommand{\pvec}{\mathbf p}\) \(\newcommand{\qvec}{\mathbf q}\) \(\newcommand{\svec}{\mathbf s}\) \(\newcommand{\tvec}{\mathbf t}\) \(\newcommand{\uvec}{\mathbf u}\) \(\newcommand{\vvec}{\mathbf v}\) \(\newcommand{\wvec}{\mathbf w}\) \(\newcommand{\xvec}{\mathbf x}\) \(\newcommand{\yvec}{\mathbf y}\) \(\newcommand{\zvec}{\mathbf z}\) \(\newcommand{\rvec}{\mathbf r}\) \(\newcommand{\mvec}{\mathbf m}\) \(\newcommand{\zerovec}{\mathbf 0}\) \(\newcommand{\onevec}{\mathbf 1}\) \(\newcommand{\real}{\mathbb R}\) \(\newcommand{\twovec}[2]{\left[\begin{array}{r}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\ctwovec}[2]{\left[\begin{array}{c}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\threevec}[3]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\cthreevec}[3]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\fourvec}[4]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\cfourvec}[4]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\fivevec}[5]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\cfivevec}[5]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\mattwo}[4]{\left[\begin{array}{rr}#1 \amp #2 \\ #3 \amp #4 \\ \end{array}\right]}\) \(\newcommand{\laspan}[1]{\text{Span}\{#1\}}\) \(\newcommand{\bcal}{\cal B}\) \(\newcommand{\ccal}{\cal C}\) \(\newcommand{\scal}{\cal S}\) \(\newcommand{\wcal}{\cal W}\) \(\newcommand{\ecal}{\cal E}\) \(\newcommand{\coords}[2]{\left\{#1\right\}_{#2}}\) \(\newcommand{\gray}[1]{\color{gray}{#1}}\) \(\newcommand{\lgray}[1]{\color{lightgray}{#1}}\) \(\newcommand{\rank}{\operatorname{rank}}\) \(\newcommand{\row}{\text{Row}}\) \(\newcommand{\col}{\text{Col}}\) \(\renewcommand{\row}{\text{Row}}\) \(\newcommand{\nul}{\text{Nul}}\) \(\newcommand{\var}{\text{Var}}\) \(\newcommand{\corr}{\text{corr}}\) \(\newcommand{\len}[1]{\left|#1\right|}\) \(\newcommand{\bbar}{\overline{\bvec}}\) \(\newcommand{\bhat}{\widehat{\bvec}}\) \(\newcommand{\bperp}{\bvec^\perp}\) \(\newcommand{\xhat}{\widehat{\xvec}}\) \(\newcommand{\vhat}{\widehat{\vvec}}\) \(\newcommand{\uhat}{\widehat{\uvec}}\) \(\newcommand{\what}{\widehat{\wvec}}\) \(\newcommand{\Sighat}{\widehat{\Sigma}}\) \(\newcommand{\lt}{<}\) \(\newcommand{\gt}{>}\) \(\newcommand{\amp}{&}\) \(\definecolor{fillinmathshade}{gray}{0.9}\)
- Define a discrete probability distribution as a list of all possible outcomes of a random variable, along with their associated probabilities.
- Use discrete probability distributions to model real-world scenarios such as dice rolls or item counts.
- Calculate key measures from the distribution, including mean, variance, standard deviation, and expected value.
As a reminder, a variable, or what will be called the random variable from now on, is represented by the letter x, which means a quantitative (numerical) variable that is measured or observed in an experiment.
Also, remember that there are different types of quantitative variables, called discrete or continuous. What is the difference between discrete and continuous data? Discrete data can only take on particular values in a range, while continuous data can take on any value in a range. Discrete data usually arises from counting, while continuous data usually arises from measuring.
Probability Distribution
How tall is a plant given a new fertilizer? Continuous. This is something you measure. How many fleas are on prairie dogs in a colony? Discrete. The number of fleas can be counted.
A random variable has outcomes that are determined by chance, and it is what is being measured. A discrete random variable is a random variable whose outcomes are countable whole-number values. For example, in the problem of "how tall is a plant given a new fertilizer?", the random variable is the height of the plant given a new fertilizer. For "how many fleas are on prairie dogs in a colony," the discrete random variable is the number of fleas on a prairie dog in a colony.
Suppose all the outcomes of a discrete random variable are listed together with their corresponding probabilities. A distribution is formed, but now it is called a probability distribution since it involves probabilities. A probability distribution is an assignment of probabilities to the values of the random variable. Furthermore, the sum of all the probabilities in the distribution must equal 1.
Note: The abbreviation of pdf is used for a probability distribution function.
For probability distributions, \(0 \leq P(x) \leq 1 \operatorname{and} \sum P(x)=1\)
The 2010 U.S. Census found the chance of a household being a certain size. Is the variable X a discrete random variable? Also, determine whether the distribution represents a probability distribution. If it does not, explain why. The data is in Example \(\PageIndex{1}\) ("Households by age," 2013).
| Size of household | 1 | 2 | 3 | 4 | 5 | 6 | 7 or more |
|---|---|---|---|---|---|---|---|
| Probability | 26.7% | 33.6% | 15.8% | 13.7% | 6.3% | 2.4% | 1.5% |
Solution
In this case, the random variable is x = the number of people in a household. Since you are counting the number of people in a household, this is a discrete random variable.
This is a probability distribution since you have the x value and the probabilities that go with it. All of the probabilities are between zero and one, and the sum of all is one.
You can give a probability distribution in table form (as in the Example above) or as a graph. The graph looks like a histogram. A probability distribution is a relative frequency distribution based on a very large sample.
The 2010 U.S. Census found the chance of a household being a certain size. The data is in the table ("Households by age," 2013). Draw a histogram of the probability distribution.
| Size of household | 1 | 2 | 3 | 4 | 5 | 6 | 7 or more |
|---|---|---|---|---|---|---|---|
| Probability | 26.7% | 33.6% | 15.8% | 13.7% | 6.3% | 2.4% | 1.5% |
Solution
State random variable:
x = number of people in a household
You draw a histogram, with the x values on the horizontal axis representing the classes (for the 7 or more categories, just call it 7) and the probabilities on the vertical axis representing the probabilities.
Notice this graph is skewed right.
Just as with any data set, the mean and standard deviation can be computed. In problems involving a probability distribution function (PDF), the probability distribution is considered the population, even though the PDF, in most cases, comes from repeating an experiment many times. This is because you are using the data from repeated experiments to estimate the true probability. Since a PDF is a population, the mean and standard deviation that are calculated are the population parameters and not the sample statistics. The notation used is the same as the notation for population mean and population standard deviation that was used in Chapter 3.
Expected Value, Variance, and Standard Deviation
The mean can be thought of as the expected value. It is the value you expect to get if the trials are repeated an infinite number of times. The mean or expected value does not need to be a whole number, even if the possible values of x are whole numbers.
For a discrete probability distribution function,
The mean or expected value is \(\mu=\sum x P(x)\)
The variance is \(\sigma^2=\left\lbrack\Sigma x\cdot p\left(x\right)\right\rbrack-\mu^2\)
The standard deviation is \(\sigma=\sqrt{\left\lbrack\Sigma x\cdot p\left(x\right)\right\rbrack-\mu^2}\)
where x = the value of the random variable and P(x) = the probability corresponding to a particular x value.
The 2010 U.S. Census found the chance of a household being a certain size. The data is in the table ("Households by age," 2013).
| Size of household | 1 | 2 | 3 | 4 | 5 | 6 | 7 or more |
|---|---|---|---|---|---|---|---|
| Probability | 26.7% | 33.6% | 15.8% | 13.7% | 6.3% | 2.4% | 1.5% |
- Find the mean
- Find the variance
- Find the standard deviation
- Use a TI-83/84 to calculate the mean and standard deviation
Solution
State random variable:
x= number of people in a household
a. To find the mean, it is easier just to use a table as shown below. Consider the category 7 or more to be just 7. The formula for the mean says to multiply the x value by the P(x) value, so add a row to the table for this calculation. Also, convert all P(x) to decimal form.
| \(X\) | \(P(X)\) | \(X \cdot P(X)\) |
|---|---|---|
| 1 | 0.267 | 0.267 |
| 2 | 0.336 | 0.672 |
| 3 | 0.158 | 0.474 |
| 4 | 0.137 | 0.548 |
| 5 | 0.063 | 0.315 |
| 6 | 0.024 | 0.144 |
| 7 | 0.015 | 0.098 |
| No sum needed | \(\Sigma = 1\) | \(\Sigma X \cdot P(X) = 2.525\) |
This is the mean or the expected value is the sum of the last column, which is \(\mu\) = 2.525 people. This means that you expect a household in the U.S. to have 2.525 people in it. Now, of course, you can’t have half a person, but what this tells you is that you expect a household to have either 2 or 3 people, with a little more 3-person households than 2-person households.
b. To find the variance, again, it is easier to use a table version than to try to use just the formula in a line. Looking at the formula, you will notice that the first operation that you should do is to subtract the mean from each x value. Then, you square each of these values. Then, you multiply each of these answers by the probability of each x value. Finally, you add up all of these values.
| \(X\) | \(P(X)\) | \(X^2 \cdot P(X)\) |
|---|---|---|
| 1 | 0.267 | 0.267 |
| 2 | 0.336 | 1.344 |
| 3 | 0.158 | 1.422 |
| 4 | 0.137 | 2.192 |
| 5 | 0.063 | 1.575 |
| 6 | 0.024 | 0.864 |
| 7 | 0.015 | 0.735 |
| No sum needed | \(\Sigma = 1\) | \(\Sigma X^2 \cdot P(X) = 8.399\) |
To find the variance, add the values in the last column and subtract the square of the mean from the sum. It is \(\sigma^{2}=8.399 - 2.525^{2} = 8.399- 6.376 = 2.023. (Note: Try not to round your numbers too much so you aren’t creating rounding errors in your answer. The numbers in the table above were rounded off because of space limitations, but the answer was calculated using many decimal places.)
c. To find the standard deviation, just take the square root of the variance, \(\sigma=\sqrt{2.023375} \approx 1.422454\) people. This means that you can expect a U.S. household to have 2.525 people in it, with a standard deviation of 1.42 people.
d. Go into the [STAT] menu, then the [Edit] menu. Type the x values into [L1] and the P(x) values into [L2]. Then go into the [STAT] menu, then the [CALC] menu. Choose [1:1-Var Stats]. This will put 1-Var Stats on the home screen. Now type in [L1], [L2] (there is a comma between [L1] and [L2]) and then press [ENTER]. If you have the newer operating system on the TI-84, then your input will be slightly different. You will see the output in the figure below.
The mean is 2.525 people, and the standard deviation is 1.422 people.
To compute the variance from the calculator output. Type the standard deviation, including all digits, into the calculator. Then select the square function and enter. Round the answer to three decimals. Thus, the variance is 2.023.
Exercises
A political analyst studies how many major policy reforms a legislative body passes in a single session. Based on historical data, the probability distribution is provided below. Perform the following steps.
- Draw a probability distribution graph (bar chart) with:
- X-axis: Number of policies supported
- Y-axis: Probability
- Describe the shape of the distribution (e.g., symmetric, skewed left, skewed right).
- Identify the most likely outcome and explain what it means in context.
- Based on the graph, are voters more likely to support fewer policies (0–1) or more policies (3–4)? Justify your answer using probabilities.
- What does this distribution suggest about overall voter behavior in this election?
| Number of Reforms (X) | 0 | 1 | 2 | 3 | 4 |
|---|---|---|---|---|---|
| P(X) | 0.10 | 0.25 | 0.30 | 0.20 | 0.15 |
- Answer
-
- See the graph below
Figure \(\PageIndex{3}\): The image shows a bar graph representing the probability distribution of the number of reforms. The horizontal axis is labeled “Number of Reforms (X)” and displays values from 0 to 4, while the vertical axis is labeled “Probability P(X).” Each bar corresponds to a specific number of reforms and its associated probability: 0 reforms (0.10), 1 reform (0.25), 2 reforms (0.30), 3 reforms (0.20), and 4 reforms (0.15). The bars are evenly spaced and labeled with their probability values, making it easy to compare the likelihood of each outcome. - The distribution is slightly skewed to the right.
- The most likely outcome is: X = 2 with probability 0.35. Voters are most likely to support 2 policies in this election.
- Fewer policies (0–1):P(0)+P(1)=0.05+0.20=0.25 and more policies (3–4): P(3)+P(4)=0.25+0.15=0.40. Voters are more likely to support more policies (3–4) since 0.4 > 0.25.
- This distribution suggests: Voters tend to support a moderate number of policies (around 2–3), there is very low support for 0 policies, and it is rare. Overall, there is a slight tendency toward supporting multiple policies rather than very few.
A college is studying the number of philosophy majors in randomly selected student discussion groups, where X represents the number of philosophy majors in a group. The probability that there are 0 philosophy majors is 0.15, the probability that there is 1 philosophy major is 0.30, the probability that there are 2 philosophy majors is 0.25, the probability that there are 3 philosophy majors is 0.20, and the probability that there are 4 philosophy majors is 0.10. Using this information, answer the following questions.
- Construct a discrete probability distribution table.
- Interpret the meaning of P(2).
- Draw a bar graph of the probability distribution
- Determine which outcome is most likely based on your graph.
- Answer
-
- The table is provided below.
Probability Distribution of Number of Philosophy Majors Number of Philosophy Majors (X) 0 1 2 3 4 P(X) 0.15 0.30 0.25 0.20 0.10 - There is a 25% chance that a randomly selected discussion group will have exactly 2 philosophy majors.
- The graph is provided below.
Figure \(\PageIndex{4}\): The image shows a bar graph representing the probability distribution of the number of philosophy majors in a group. The horizontal axis is labeled “Number of Philosophy Majors (X)” and includes values from 0 to 4, while the vertical axis is labeled “P(X).” Each bar is colored purple and represents the probability associated with each value of X. The tallest bar occurs at 2 philosophy majors, indicating it is the most likely outcome, while shorter bars appear at 0 and 4, showing lower probabilities. The graph provides a clear visual comparison of how likely each number of philosophy majors is within a group.- The most likely outcome is 1 philosophy major per group.
A researcher is studying how many new public policies a city council adopts in a year. The probability distribution is given below, and then answer the questions below the table.
| Number of Reforms (X) | 0 | 1 | 2 | 3 | 4 |
|---|---|---|---|---|---|
| P(X) | 0.10 | 0.25 | 0.30 | 0.20 | 0.15 |
- Find the mean of the distribution.
- Find the variance of the distribution.
- Find the standard deviation.
- Interpret the meaning of the mean and standard deviation in context.
- Answer
-
- \(\mu = 2.15\)
- \(\sigma^2 = 1.43\)
- \(\sigma = 1.19\)
- The mean suggests that, on average, about 2 policies are adopted per year. The standard deviation indicates a moderate spread, meaning the number of policies typically varies by about 1 policy from the mean.
-
A music artist schedules a varying number of live shows each month. Based on past data, the probability distribution for the number of shows is given below, and answer the questions below it.
| Number of Shows (X) | 1 | 2 | 3 | 4 | 5 |
|---|---|---|---|---|---|
| P(X) | 0.15 | 0.25 | 0.30 | 0.20 | 0.10 |
- Find the mean.
- Find the variance.
- Find the standard deviation.
- Interpret the results in the context of the artist’s performance schedule.
- Answer
-
- \(\mu = 2.85\)
- \(\sigma^2 = 1.43\)
- \(\sigma = 1.19\)
- The artist performs about 2.85 shows per month on average (roughly 3 shows). The standard deviation of 1.19 indicates that the number of shows typically varies by about 1 show from the average. This suggests a fairly consistent performance schedule, with occasional busier or lighter months.
Authors
"5.1: Basics of Probability Distributions"(opens in new window) by Toros Berberyan(opens in new window), Tracy Nguyen(opens in new window), and Alfie Swan(opens in new window) is licensed under CC BY-SA 4.0(opens in new window)
Attributions
"5.1: Basics of Probability Distributions"(opens in new window) by Kathryn Kozak(opens in new window) is licensed under CC BY-SA 4.0(opens in new window)


