Lecture 6 - Expectation and Moments of a Random Variable
In the previous lectures we introduced random variables and several important probability distributions. We have seen how a probability distribution describes the possible values of a random variable and assigns probabilities to them.
It is often useful, however, to summarize a probability distribution by a small number of numerical quantities.
For example:
What is the typical value of a random variable?
How much does it fluctuate around its typical value?
How can we quantify its spread?
How can we describe asymmetry or other features of a distribution?
The most important quantities used for this purpose are the expectation , the moments , and the variance .
In this lecture we develop these concepts for a single random variable. We also introduce generating functions for discrete random variables, which provide a convenient way of computing moments and studying sums of independent random variables.
The expected value of a discrete random variable ¶ Let ξ \xi ξ be a discrete random variable with probability mass function P ξ P_\xi P ξ .
The expected value of ξ \xi ξ is defined by
E ξ = ∑ x x P ξ ( x ) , \mathbb{E}\xi
=
\sum_x xP_\xi(x), E ξ = x ∑ x P ξ ( x ) , provided that the series is absolutely convergent.
The expectation can be interpreted as the long-run average value obtained by repeatedly observing the random experiment.
If the possible values of ξ \xi ξ are
x 1 , x 2 , … x_1,x_2,\ldots x 1 , x 2 , … with probabilities
p 1 , p 2 , … , p_1,p_2,\ldots, p 1 , p 2 , … , then
E ξ = x 1 p 1 + x 2 p 2 + ⋯ . \mathbb{E}\xi
=
x_1p_1+x_2p_2+\cdots. E ξ = x 1 p 1 + x 2 p 2 + ⋯ . Thus, the expectation is a probability-weighted average of the possible values.
Let ξ \xi ξ be the outcome of a fair six-sided die. Then
P ξ ( k ) = 1 6 , k = 1 , … , 6. P_\xi(k)=\frac16,
\qquad
k=1,\ldots,6. P ξ ( k ) = 6 1 , k = 1 , … , 6. Therefore,
E ξ = ∑ k = 1 6 k 1 6 = 1 + 2 + 3 + 4 + 5 + 6 6 = 7 2 . \begin{aligned}
\mathbb{E}\xi
&=
\sum_{k=1}^6 k\frac16\\
&=
\frac{1+2+3+4+5+6}{6}\\
&=
\frac72.
\end{aligned} E ξ = k = 1 ∑ 6 k 6 1 = 6 1 + 2 + 3 + 4 + 5 + 6 = 2 7 . Thus the expected value is
E ξ = 3.5. \mathbb{E}\xi=3.5. E ξ = 3.5. Notice that 3.5 is not a possible outcome of the die. The expected value need not itself be a value that the random variable can actually take.
Expectation of a function of a random variable ¶ Suppose that
η = φ ( ξ ) \eta=\varphi(\xi) η = φ ( ξ ) for some function φ \varphi φ .
There is no need to determine the probability distribution of η \eta η before computing its expectation. Instead,
E φ ( ξ ) = ∑ x φ ( x ) P ξ ( x ) , \mathbb{E}\varphi(\xi)
=
\sum_x \varphi(x)P_\xi(x), E φ ( ξ ) = x ∑ φ ( x ) P ξ ( x ) , provided that the sum is well defined.
This result is sometimes called the law of the unconscious statistician .
For example, if ξ \xi ξ is the outcome of a fair die and we want the expected value of ξ 2 \xi^2 ξ 2 , then
E ξ 2 = ∑ k = 1 6 k 2 1 6 = 1 + 4 + 9 + 16 + 25 + 36 6 = 91 6 . \begin{aligned}
\mathbb{E}\xi^2
&=
\sum_{k=1}^6 k^2\frac16\\
&=
\frac{1+4+9+16+25+36}{6}\\
&=
\frac{91}{6}.
\end{aligned} E ξ 2 = k = 1 ∑ 6 k 2 6 1 = 6 1 + 4 + 9 + 16 + 25 + 36 = 6 91 . We will see that this quantity is closely related to the variance of ξ \xi ξ .
Expectation of a continuous random variable ¶ For a continuous random variable with density p ξ p_\xi p ξ , the corresponding definition is obtained by replacing the sum with an integral:
E ξ = ∫ − ∞ + ∞ x p ξ ( x ) d x , \mathbb{E}\xi
=
\int_{-\infty}^{+\infty}
x p_\xi(x)\,dx, E ξ = ∫ − ∞ + ∞ x p ξ ( x ) d x , provided that the integral is absolutely convergent.
More generally,
E φ ( ξ ) = ∫ − ∞ + ∞ φ ( x ) p ξ ( x ) d x . \mathbb{E}\varphi(\xi)
=
\int_{-\infty}^{+\infty}
\varphi(x)p_\xi(x)\,dx. E φ ( ξ ) = ∫ − ∞ + ∞ φ ( x ) p ξ ( x ) d x . Thus, the discrete and continuous cases have exactly the same structure:
E φ ( ξ ) = { ∑ x φ ( x ) P ξ ( x ) , discrete , ∫ − ∞ + ∞ φ ( x ) p ξ ( x ) d x , continuous . \boxed{
\mathbb{E}\varphi(\xi)
=
\begin{cases}
\displaystyle
\sum_x\varphi(x)P_\xi(x),
& \text{discrete},\\[12pt]
\displaystyle
\int_{-\infty}^{+\infty}
\varphi(x)p_\xi(x)\,dx,
& \text{continuous}.
\end{cases}} E φ ( ξ ) = ⎩ ⎨ ⎧ x ∑ φ ( x ) P ξ ( x ) , ∫ − ∞ + ∞ φ ( x ) p ξ ( x ) d x , discrete , continuous . Properties of expectation ¶ Expectation satisfies several fundamental properties.
Constant ¶ If c c c is a constant, then
Scaling ¶ For any constant c c c ,
E ( c ξ ) = c E ξ . \mathbb{E}(c\xi)
=
c\mathbb{E}\xi. E ( c ξ ) = c E ξ . Linearity ¶ For random variables ξ \xi ξ and η \eta η for which the expectations exist,
E ( ξ + η ) = E ξ + E η . \mathbb{E}(\xi+\eta)
=
\mathbb{E}\xi+\mathbb{E}\eta. E ( ξ + η ) = E ξ + E η . More generally,
E ( c 1 ξ + c 2 η ) = c 1 E ξ + c 2 E η . \mathbb{E}(c_1\xi+c_2\eta)
=
c_1\mathbb{E}\xi+c_2\mathbb{E}\eta. E ( c 1 ξ + c 2 η ) = c 1 E ξ + c 2 E η . Importantly, independence is not required for linearity of expectation .
Monotonicity ¶ If
with probability one, then
E ξ ≤ E η . \mathbb{E}\xi\leq\mathbb{E}\eta. E ξ ≤ E η . Absolute value ¶ The triangle inequality for expectation gives
∣ E ξ ∣ ≤ E ∣ ξ ∣ . |\mathbb{E}\xi|
\leq
\mathbb{E}|\xi|. ∣ E ξ ∣ ≤ E ∣ ξ ∣. This inequality is useful when establishing the existence of expectations and moments.
Expectation of a Bernoulli random variable ¶ Let
ξ ∼ Bern ( p ) . \xi\sim\operatorname{Bern}(p). ξ ∼ Bern ( p ) . Thus
P ξ ( 0 ) = 1 − p , P ξ ( 1 ) = p . P_\xi(0)=1-p,
\qquad
P_\xi(1)=p. P ξ ( 0 ) = 1 − p , P ξ ( 1 ) = p . Its expectation is
E ξ = 0 ( 1 − p ) + 1 p = p . \begin{aligned}
\mathbb{E}\xi
&=
0(1-p)+1p\\
&=p.
\end{aligned} E ξ = 0 ( 1 − p ) + 1 p = p . Therefore,
E ξ = p . \boxed{\mathbb{E}\xi=p.} E ξ = p . This simple result has an important interpretation: the expected value of an indicator variable is the probability of the event it represents.
If A A A is an event, define its indicator by
1 A = { 1 , A occurs , 0 , A does not occur . \mathbf{1}_A
=
\begin{cases}
1, & A\text{ occurs},\\
0, & A\text{ does not occur}.
\end{cases} 1 A = { 1 , 0 , A occurs , A does not occur . Then
E 1 A = P ( A ) . \mathbb{E}\mathbf{1}_A
=
\mathbb{P}(A). E 1 A = P ( A ) . Indicator variables will be particularly useful when studying sums of random variables.
Expectation of the binomial distribution ¶ Recall that
S n = ξ 1 + ⋯ + ξ n , S_n=\xi_1+\cdots+\xi_n, S n = ξ 1 + ⋯ + ξ n , where the ξ k \xi_k ξ k are independent Bernoulli random variables with parameter p p p .
Since
E ξ k = p , \mathbb{E}\xi_k=p, E ξ k = p , linearity of expectation gives
E S n = E ( ξ 1 + ⋯ + ξ n ) = ∑ k = 1 n E ξ k = n p . \begin{aligned}
\mathbb{E}S_n
&=
\mathbb{E}(\xi_1+\cdots+\xi_n)\\
&=
\sum_{k=1}^n\mathbb{E}\xi_k\\
&=
np.
\end{aligned} E S n = E ( ξ 1 + ⋯ + ξ n ) = k = 1 ∑ n E ξ k = n p . Hence, if
S n ∼ Bin ( n , p ) , S_n\sim\operatorname{Bin}(n,p), S n ∼ Bin ( n , p ) , then
E S n = n p . \boxed{\mathbb{E}S_n=np.} E S n = n p . The result does not require us to perform the binomial sum directly.
The interpretation is natural: if each of n n n trials has success probability p p p , then the expected number of successes is n p np n p .
Expectation of the Normal distribution ¶ Let
ξ ∼ N ( μ , σ 2 ) . \xi \sim \mathcal{N}(\mu, \sigma^2). ξ ∼ N ( μ , σ 2 ) . The probability density function is
f ξ ( x ) = 1 2 π σ 2 e − ( x − μ ) 2 2 σ 2 , x ∈ R . f_\xi(x) = \frac{1}{\sqrt{2\pi\sigma^2}} e^{-\frac{(x-\mu)^2}{2\sigma^2}}, \qquad x \in \mathbb{R}. f ξ ( x ) = 2 π σ 2 1 e − 2 σ 2 ( x − μ ) 2 , x ∈ R . The expectation is
E ξ = ∫ − ∞ ∞ x f ξ ( x ) d x = ∫ − ∞ ∞ x 1 2 π σ 2 e − ( x − μ ) 2 2 σ 2 d x = ∫ − ∞ ∞ ( y σ + μ ) 1 2 π σ 2 e − y 2 2 σ d y ( substituting y = x − μ σ , d x = σ d y ) = σ ∫ − ∞ ∞ y 1 2 π e − y 2 2 d y + μ ∫ − ∞ ∞ 1 2 π e − y 2 2 d y = σ ⋅ 0 + μ ⋅ 1 = μ . \begin{split}
\mathbb{E}\xi = & \int_{-\infty}^\infty x f_\xi(x) \, dx \\
= & \int_{-\infty}^\infty x \frac{1}{\sqrt{2\pi\sigma^2}} e^{-\frac{(x-\mu)^2}{2\sigma^2}} \, dx \\
= & \int_{-\infty}^\infty (y\sigma + \mu) \frac{1}{\sqrt{2\pi\sigma^2}} e^{-\frac{y^2}{2}} \sigma \, dy \quad \left(\text{substituting } y = \frac{x-\mu}{\sigma}, \, dx = \sigma \, dy\right) \\
= & \sigma \int_{-\infty}^\infty y \frac{1}{\sqrt{2\pi}} e^{-\frac{y^2}{2}} \, dy + \mu \int_{-\infty}^\infty \frac{1}{\sqrt{2\pi}} e^{-\frac{y^2}{2}} \, dy \\
= & \sigma \cdot 0 + \mu \cdot 1 \\
= & \mu.
\end{split} E ξ = = = = = = ∫ − ∞ ∞ x f ξ ( x ) d x ∫ − ∞ ∞ x 2 π σ 2 1 e − 2 σ 2 ( x − μ ) 2 d x ∫ − ∞ ∞ ( y σ + μ ) 2 π σ 2 1 e − 2 y 2 σ d y ( substituting y = σ x − μ , d x = σ d y ) σ ∫ − ∞ ∞ y 2 π 1 e − 2 y 2 d y + μ ∫ − ∞ ∞ 2 π 1 e − 2 y 2 d y σ ⋅ 0 + μ ⋅ 1 μ . Thus,
E ξ = μ . \mathbb{E}\xi = \mu. E ξ = μ . Moments ¶ The k k k -th raw moment of a random variable ξ \xi ξ is the expectation
m k = E ξ k , m_k
=
\mathbb{E}\xi^k, m k = E ξ k , whenever the expectation exists.
Thus,
m 1 = E ξ , m 2 = E ξ 2 , m 3 = E ξ 3 , m_1=\mathbb{E}\xi,
\qquad
m_2=\mathbb{E}\xi^2,
\qquad
m_3=\mathbb{E}\xi^3, m 1 = E ξ , m 2 = E ξ 2 , m 3 = E ξ 3 , and so on.
The first moment is the expected value.
Higher moments provide additional information about the distribution.
For example, the second moment E ξ 2 \mathbb{E}\xi^2 E ξ 2 is used to quantify the spread of a random variable, while higher moments can be used to characterize features such as asymmetry and tail behavior.
It is often useful to consider centered moments :
μ k = E ( ξ − E ξ ) k . \mu_k
=
\mathbb{E}(\xi-\mathbb{E}\xi)^k. μ k = E ( ξ − E ξ ) k . The first centered moment is always zero:
E ( ξ − E ξ ) = 0. \mathbb{E}(\xi-\mathbb{E}\xi)=0. E ( ξ − E ξ ) = 0. The second centered moment is the variance.
Variance ¶ The variance of a random variable ξ \xi ξ is defined by
Var ( ξ ) = E [ ( ξ − E ξ ) 2 ] . \operatorname{Var}(\xi)
=
\mathbb{E}
\left[
(\xi-\mathbb{E}\xi)^2
\right]. Var ( ξ ) = E [ ( ξ − E ξ ) 2 ] . It measures the average squared distance of ξ \xi ξ from its expected value.
The variance is always nonnegative:
Var ( ξ ) ≥ 0. \operatorname{Var}(\xi)\geq0. Var ( ξ ) ≥ 0. Indeed, the random variable
( ξ − E ξ ) 2 (\xi-\mathbb{E}\xi)^2 ( ξ − E ξ ) 2 is nonnegative.
A useful alternative expression follows by expanding the square:
Var ( ξ ) = E [ ξ 2 − 2 ξ E ξ + ( E ξ ) 2 ] = E ξ 2 − 2 ( E ξ ) 2 + ( E ξ ) 2 . \begin{aligned}
\operatorname{Var}(\xi)
&=
\mathbb{E}
\left[
\xi^2-2\xi\mathbb{E}\xi+(\mathbb{E}\xi)^2
\right]\\
&=
\mathbb{E}\xi^2
-2(\mathbb{E}\xi)^2
+(\mathbb{E}\xi)^2.
\end{aligned} Var ( ξ ) = E [ ξ 2 − 2 ξ E ξ + ( E ξ ) 2 ] = E ξ 2 − 2 ( E ξ ) 2 + ( E ξ ) 2 . Therefore,
Var ( ξ ) = E ξ 2 − ( E ξ ) 2 . \boxed{
\operatorname{Var}(\xi)
=
\mathbb{E}\xi^2-(\mathbb{E}\xi)^2.
} Var ( ξ ) = E ξ 2 − ( E ξ ) 2 . This formula is often much more convenient for computations.
Variance of a Bernoulli random variable ¶ Let
ξ ∼ Bern ( p ) . \xi\sim\operatorname{Bern}(p). ξ ∼ Bern ( p ) . Since ξ \xi ξ takes only the values 0 and 1,
Therefore,
E ξ 2 = E ξ = p . \mathbb{E}\xi^2=\mathbb{E}\xi=p. E ξ 2 = E ξ = p . Using (18) ,
Var ( ξ ) = p − p 2 = p ( 1 − p ) . \begin{aligned}
\operatorname{Var}(\xi)
&=
p-p^2\\
&=
p(1-p).
\end{aligned} Var ( ξ ) = p − p 2 = p ( 1 − p ) . Hence,
Var ( ξ ) = p ( 1 − p ) . \boxed{
\operatorname{Var}(\xi)=p(1-p).
} Var ( ξ ) = p ( 1 − p ) . The variance is largest when p = 1 / 2 p=1/2 p = 1/2 and becomes small when p p p is close to either 0 or 1.
Variance of a sum of independent random variables ¶ Let ξ \xi ξ and η \eta η be independent random variables with finite second moments.
We begin with
Var ( ξ + η ) = E [ ( ξ + η − E ξ − E η ) 2 ] . \operatorname{Var}(\xi+\eta)
=
\mathbb{E}
\left[
(\xi+\eta-\mathbb{E}\xi-\mathbb{E}\eta)^2
\right]. Var ( ξ + η ) = E [ ( ξ + η − E ξ − E η ) 2 ] . Expanding the square,
Var ( ξ + η ) = E ( ξ − E ξ ) 2 + E ( η − E η ) 2 + 2 E [ ( ξ − E ξ ) ( η − E η ) ] . \begin{aligned}
\operatorname{Var}(\xi+\eta)
&=
\mathbb{E}(\xi-\mathbb{E}\xi)^2
+
\mathbb{E}(\eta-\mathbb{E}\eta)^2\\
&\quad+
2\mathbb{E}
\left[
(\xi-\mathbb{E}\xi)
(\eta-\mathbb{E}\eta)
\right].
\end{aligned} Var ( ξ + η ) = E ( ξ − E ξ ) 2 + E ( η − E η ) 2 + 2 E [ ( ξ − E ξ ) ( η − E η ) ] . Independence implies
E [ ( ξ − E ξ ) ( η − E η ) ] = E ( ξ − E ξ ) E ( η − E η ) = 0. \mathbb{E}
\left[
(\xi-\mathbb{E}\xi)
(\eta-\mathbb{E}\eta)
\right]
=
\mathbb{E}(\xi-\mathbb{E}\xi)
\mathbb{E}(\eta-\mathbb{E}\eta)
=
0. E [ ( ξ − E ξ ) ( η − E η ) ] = E ( ξ − E ξ ) E ( η − E η ) = 0. Consequently,
Var ( ξ + η ) = Var ( ξ ) + Var ( η ) . \boxed{
\operatorname{Var}(\xi+\eta)
=
\operatorname{Var}(\xi)
+
\operatorname{Var}(\eta).
} Var ( ξ + η ) = Var ( ξ ) + Var ( η ) . By induction, if ξ 1 , … , ξ n \xi_1,\ldots,\xi_n ξ 1 , … , ξ n are independent,
Var ( ∑ k = 1 n ξ k ) = ∑ k = 1 n Var ( ξ k ) . \boxed{
\operatorname{Var}
\left(
\sum_{k=1}^n\xi_k
\right)
=
\sum_{k=1}^n\operatorname{Var}(\xi_k).
} Var ( k = 1 ∑ n ξ k ) = k = 1 ∑ n Var ( ξ k ) . Variance of the binomial distribution ¶ For
S n = ξ 1 + ⋯ + ξ n , S_n=\xi_1+\cdots+\xi_n, S n = ξ 1 + ⋯ + ξ n , with independent Bernoulli random variables,
Var ( ξ k ) = p ( 1 − p ) . \operatorname{Var}(\xi_k)=p(1-p). Var ( ξ k ) = p ( 1 − p ) . Therefore, (21) gives
Var ( S n ) = ∑ k = 1 n p ( 1 − p ) = n p ( 1 − p ) . \begin{aligned}
\operatorname{Var}(S_n)
&=
\sum_{k=1}^n p(1-p)\\
&=
np(1-p).
\end{aligned} Var ( S n ) = k = 1 ∑ n p ( 1 − p ) = n p ( 1 − p ) . Thus,
Var ( S n ) = n p ( 1 − p ) . \boxed{
\operatorname{Var}(S_n)=np(1-p).
} Var ( S n ) = n p ( 1 − p ) . We have therefore obtained both parameters that appeared in the normal approximation of Lecture 5:
E S n = n p , Var ( S n ) = n p ( 1 − p ) . \mathbb{E}S_n=np,
\qquad
\operatorname{Var}(S_n)=np(1-p). E S n = n p , Var ( S n ) = n p ( 1 − p ) . Variance of the Normal distribution ¶ Let ξ ∼ N ( μ , σ 2 ) \xi \sim \mathcal{N}(\mu, \sigma^2) ξ ∼ N ( μ , σ 2 ) , as for the other cases it is useful to use the relation
Var ( ξ ) = E ( ξ − μ ) 2 = E ξ 2 − ( E ξ ) 2 . \operatorname{Var}(\xi) = \mathbb{E}(\xi - \mu)^2 = \mathbb{E}\xi^2 - (\mathbb{E}\xi)^2. Var ( ξ ) = E ( ξ − μ ) 2 = E ξ 2 − ( E ξ ) 2 . where we have used the fact that E ξ = μ E\xi = \mu E ξ = μ for the Normal distribution.
We justt to compute the second order moment E ξ 2 \mathbb{E}\xi^2 E ξ 2 as
E ξ 2 = ∫ − ∞ ∞ x 2 1 2 π σ 2 e − ( x − μ ) 2 2 σ 2 d x = ∫ − ∞ ∞ ( y σ + μ ) 2 1 2 π e − y 2 2 d y ( substituting y = x − μ σ , d x = σ d y ) = ∫ − ∞ ∞ ( y 2 σ 2 + 2 μ σ y + μ 2 ) 1 2 π e − y 2 2 d y = σ 2 ∫ − ∞ ∞ y 2 1 2 π e − y 2 2 d y + 2 μ σ ∫ − ∞ ∞ y 1 2 π e − y 2 2 d y + μ 2 ∫ − ∞ ∞ 1 2 π e − y 2 2 d y = σ 2 ( [ − y 1 2 π e − y 2 2 ] − ∞ ∞ − ∫ − ∞ ∞ ( − 1 2 π e − y 2 2 ) d y ) + 2 μ σ ⋅ 0 + μ 2 ⋅ 1 = σ 2 ( 0 + ∫ − ∞ ∞ 1 2 π e − y 2 2 d y ) + μ 2 = σ 2 ⋅ 1 + μ 2 = μ 2 + σ 2 . \begin{split}
\mathbb{E}\xi^2 = & \int_{-\infty}^\infty x^2 \frac{1}{\sqrt{2\pi\sigma^2}} e^{-\frac{(x-\mu)^2}{2\sigma^2}} \, dx \\
= & \int_{-\infty}^\infty (y\sigma + \mu)^2 \frac{1}{\sqrt{2\pi}} e^{-\frac{y^2}{2}} \, dy \quad \left(\text{substituting } y = \frac{x-\mu}{\sigma}, \, dx = \sigma \, dy\right) \\
= & \int_{-\infty}^\infty (y^2 \sigma^2 + 2\mu\sigma y + \mu^2) \frac{1}{\sqrt{2\pi}} e^{-\frac{y^2}{2}} \, dy \\
= & \sigma^2 \int_{-\infty}^\infty y^2 \frac{1}{\sqrt{2\pi}} e^{-\frac{y^2}{2}} \, dy + 2\mu\sigma \int_{-\infty}^\infty y \frac{1}{\sqrt{2\pi}} e^{-\frac{y^2}{2}} \, dy + \mu^2 \int_{-\infty}^\infty \frac{1}{\sqrt{2\pi}} e^{-\frac{y^2}{2}} \, dy \\
= & \sigma^2 \left( \left[ -y \frac{1}{\sqrt{2\pi}} e^{-\frac{y^2}{2}} \right]_{-\infty}^\infty - \int_{-\infty}^\infty \left( -\frac{1}{\sqrt{2\pi}} e^{-\frac{y^2}{2}} \right) dy \right) + 2\mu\sigma \cdot 0 + \mu^2 \cdot 1 \\
= & \sigma^2 \left( 0 + \int_{-\infty}^\infty \frac{1}{\sqrt{2\pi}} e^{-\frac{y^2}{2}} \, dy \right) + \mu^2 \\
= & \sigma^2 \cdot 1 + \mu^2 \\
= & \mu^2 + \sigma^2.
\end{split} E ξ 2 = = = = = = = = ∫ − ∞ ∞ x 2 2 π σ 2 1 e − 2 σ 2 ( x − μ ) 2 d x ∫ − ∞ ∞ ( y σ + μ ) 2 2 π 1 e − 2 y 2 d y ( substituting y = σ x − μ , d x = σ d y ) ∫ − ∞ ∞ ( y 2 σ 2 + 2 μ σ y + μ 2 ) 2 π 1 e − 2 y 2 d y σ 2 ∫ − ∞ ∞ y 2 2 π 1 e − 2 y 2 d y + 2 μ σ ∫ − ∞ ∞ y 2 π 1 e − 2 y 2 d y + μ 2 ∫ − ∞ ∞ 2 π 1 e − 2 y 2 d y σ 2 ( [ − y 2 π 1 e − 2 y 2 ] − ∞ ∞ − ∫ − ∞ ∞ ( − 2 π 1 e − 2 y 2 ) d y ) + 2 μ σ ⋅ 0 + μ 2 ⋅ 1 σ 2 ( 0 + ∫ − ∞ ∞ 2 π 1 e − 2 y 2 d y ) + μ 2 σ 2 ⋅ 1 + μ 2 μ 2 + σ 2 . Now, substituting E ξ 2 = μ 2 + σ 2 \mathbb{E}\xi^2 = \mu^2 + \sigma^2 E ξ 2 = μ 2 + σ 2 :
Var ( ξ ) = E ξ 2 − ( E ξ ) 2 = ( μ 2 + σ 2 ) − μ 2 = σ 2 . \begin{split}
\operatorname{Var}(\xi) = & \mathbb{E}\xi^2 - (\mathbb{E}\xi)^2 \\
= & (\mu^2 + \sigma^2) - \mu^2 \\
= & \sigma^2.
\end{split} Var ( ξ ) = = = E ξ 2 − ( E ξ ) 2 ( μ 2 + σ 2 ) − μ 2 σ 2 . Thus,
Var ( ξ ) = σ 2 . \operatorname{Var}(\xi) = \sigma^2. Var ( ξ ) = σ 2 . Standard deviation ¶ The variance has the units of the square of the original quantity.
For example, if ξ \xi ξ represents a length measured in meters, then Var ( ξ ) \operatorname{Var}(\xi) Var ( ξ ) is measured in square meters.
It is often more natural to measure spread in the same units as the random variable. This leads to the standard deviation :
σ ξ = Var ( ξ ) . \sigma_\xi
=
\sqrt{\operatorname{Var}(\xi)}. σ ξ = Var ( ξ ) . Thus, the standard deviation is the positive square root of the variance.
The standard deviation measures the typical scale of fluctuations around the expected value.
For a binomial random variable,
σ S n = n p ( 1 − p ) . \sigma_{S_n}
=
\sqrt{np(1-p)}. σ S n = n p ( 1 − p ) . This is precisely the scaling that appeared in the De Moivre--Laplace theorem.
Suppose that
η = c ξ + d , \eta=c\xi+d, η = c ξ + d , where c c c and d d d are constants.
Using linearity of expectation,
E η = c E ξ + d . \mathbb{E}\eta
=
c\mathbb{E}\xi+d. E η = c E ξ + d . For the variance,
Var ( η ) = Var ( c ξ + d ) = E [ c ξ + d − ( c E ξ + d ) ] 2 = E [ c ( ξ − E ξ ) ] 2 = c 2 Var ( ξ ) . \begin{aligned}
\operatorname{Var}(\eta)
&=
\operatorname{Var}(c\xi+d)\\
&=
\mathbb{E}
\left[
c\xi+d-(c\mathbb{E}\xi+d)
\right]^2\\
&=
\mathbb{E}
\left[
c(\xi-\mathbb{E}\xi)
\right]^2\\
&=
c^2\operatorname{Var}(\xi).
\end{aligned} Var ( η ) = Var ( c ξ + d ) = E [ c ξ + d − ( c E ξ + d ) ] 2 = E [ c ( ξ − E ξ ) ] 2 = c 2 Var ( ξ ) . Therefore,
Var ( c ξ + d ) = c 2 Var ( ξ ) . \boxed{
\operatorname{Var}(c\xi+d)
=
c^2\operatorname{Var}(\xi).
} Var ( c ξ + d ) = c 2 Var ( ξ ) . In particular, adding a constant changes the mean but does not change the variance, while multiplying by c c c scales the standard deviation by ∣ c ∣ |c| ∣ c ∣ .
Moments of the normal distribution ¶ Consider
ξ ∼ N ( a , σ 2 ) . \xi\sim N(a,\sigma^2). ξ ∼ N ( a , σ 2 ) . From Lecture 5, we can write
ξ = a + σ Z , \xi=a+\sigma Z, ξ = a + σ Z , where
Z ∼ N ( 0 , 1 ) . Z\sim N(0,1). Z ∼ N ( 0 , 1 ) . The standard normal distribution is symmetric about zero. Consequently,
Its variance is
Var ( Z ) = 1. \operatorname{Var}(Z)=1. Var ( Z ) = 1. Therefore, using (25) and (26) ,
E ξ = a , Var ( ξ ) = σ 2 . \boxed{
\mathbb{E}\xi=a,
\qquad
\operatorname{Var}(\xi)=\sigma^2.
} E ξ = a , Var ( ξ ) = σ 2 . This gives the probabilistic interpretation of the two parameters of the normal distribution:
a a a is the expected value;
σ 2 \sigma^2 σ 2 is the variance;
σ \sigma σ is the standard deviation.
Moments of the exponential distribution ¶ Consider
ξ ∼ Exp ( λ ) . \xi\sim\operatorname{Exp}(\lambda). ξ ∼ Exp ( λ ) . Its expectation is
E ξ = ∫ 0 ∞ x λ e − λ x d x . \begin{aligned}
\mathbb{E}\xi
&=
\int_0^\infty x\lambda e^{-\lambda x}\,dx.
\end{aligned} E ξ = ∫ 0 ∞ x λ e − λ x d x . Integration by parts gives
∫ 0 ∞ x λ e − λ x d x = [ − x e − λ x ] 0 ∞ − ∫ 0 ∞ ( − e − λ x ) d x = lim x → ∞ − x e λ x + 0 ⋅ e − λ ⋅ 0 + [ − 1 λ e − λ x ] 0 ∞ = lim x → ∞ ( − 1 λ e − λ x ) + 1 λ e − λ 0 = 1 λ . \begin{split}
\int_0^\infty x \lambda e^{-\lambda x} \, dx = & \left[ -x e^{-\lambda x} \right]_0^\infty - \int_0^\infty \left( -e^{-\lambda x} \right) dx \\
= & \lim_{x \to \infty} \frac{-x}{e^{\lambda x}} + 0 \cdot e^{-\lambda \cdot 0} + \left[ -\frac{1}{\lambda} e^{-\lambda x} \right]_0^\infty \\
= & \lim_{x \to \infty} \left( -\frac{1}{\lambda} e^{-\lambda x} \right) + \frac{1}{\lambda} e^{-\lambda 0} \\
= & \frac{1}{\lambda}.
\end{split} ∫ 0 ∞ x λ e − λ x d x = = = = [ − x e − λ x ] 0 ∞ − ∫ 0 ∞ ( − e − λ x ) d x x → ∞ lim e λ x − x + 0 ⋅ e − λ ⋅ 0 + [ − λ 1 e − λ x ] 0 ∞ x → ∞ lim ( − λ 1 e − λ x ) + λ 1 e − λ 0 λ 1 . Hence,
E ξ = 1 λ . \boxed{
\mathbb{E}\xi=\frac1\lambda.
} E ξ = λ 1 . The second moment is
E ξ 2 = ∫ 0 ∞ x 2 λ e − λ x d x = [ − x 2 e − λ x ] 0 ∞ − ∫ 0 ∞ 2 x ( − e − λ x ) d x = lim x → ∞ − x 2 e λ x + 0 2 ⋅ e − λ ⋅ 0 + 2 λ ∫ 0 ∞ x λ e − λ x d x = 0 + 0 + 2 λ ( [ − x e − λ x ] 0 ∞ − ∫ 0 ∞ ( − e − λ x ) d x ) = 2 λ ( lim x → ∞ − x e λ x + 0 ⋅ e − λ ⋅ 0 + [ − 1 λ e − λ x ] 0 ∞ ) = 2 λ ( lim x → ∞ ( − 1 λ e − λ x ) + 1 λ e − λ ⋅ 0 ) = 2 λ 2 . \begin{split}
\mathbb{E}\xi^2 = & \int_0^\infty x^2 \lambda e^{-\lambda x} \, dx \\
= & \left[ -x^2 e^{-\lambda x} \right]_0^\infty - \int_0^\infty 2x \left( -e^{-\lambda x} \right) dx \\
= & \lim_{x \to \infty} \frac{-x^2}{e^{\lambda x}} + 0^2 \cdot e^{-\lambda \cdot 0} + \frac{2}{\lambda} \int_0^\infty x \lambda e^{-\lambda x} \, dx \\
= & 0 + 0 + \frac{2}{\lambda} \left( \left[ -x e^{-\lambda x} \right]_0^\infty - \int_0^\infty \left( -e^{-\lambda x} \right) dx \right) \\
= & \frac{2}{\lambda} \left( \lim_{x \to \infty} \frac{-x}{e^{\lambda x}} + 0 \cdot e^{-\lambda \cdot 0} + \left[ -\frac{1}{\lambda} e^{-\lambda x} \right]_0^\infty \right) \\
= & \frac{2}{\lambda} \left( \lim_{x \to \infty} \left( -\frac{1}{\lambda} e^{-\lambda x} \right) + \frac{1}{\lambda} e^{-\lambda \cdot 0} \right) \\
= & \frac{2}{\lambda^2}.
\end{split} E ξ 2 = = = = = = = ∫ 0 ∞ x 2 λ e − λ x d x [ − x 2 e − λ x ] 0 ∞ − ∫ 0 ∞ 2 x ( − e − λ x ) d x x → ∞ lim e λ x − x 2 + 0 2 ⋅ e − λ ⋅ 0 + λ 2 ∫ 0 ∞ x λ e − λ x d x 0 + 0 + λ 2 ( [ − x e − λ x ] 0 ∞ − ∫ 0 ∞ ( − e − λ x ) d x ) λ 2 ( x → ∞ lim e λ x − x + 0 ⋅ e − λ ⋅ 0 + [ − λ 1 e − λ x ] 0 ∞ ) λ 2 ( x → ∞ lim ( − λ 1 e − λ x ) + λ 1 e − λ ⋅ 0 ) λ 2 2 . Therefore,
Var ( ξ ) = E ξ 2 − ( E ξ ) 2 = 2 λ 2 − 1 λ 2 = 1 λ 2 . \begin{aligned}
\operatorname{Var}(\xi)
&=
\mathbb{E}\xi^2-(\mathbb{E}\xi)^2\\
&=
\frac{2}{\lambda^2}
-
\frac{1}{\lambda^2}\\
&=
\frac{1}{\lambda^2}.
\end{aligned} Var ( ξ ) = E ξ 2 − ( E ξ ) 2 = λ 2 2 − λ 2 1 = λ 2 1 . Thus,
Var ( ξ ) = 1 λ 2 , σ ξ = 1 λ . \boxed{
\operatorname{Var}(\xi)=\frac1{\lambda^2},
\qquad
\sigma_\xi=\frac1\lambda.
} Var ( ξ ) = λ 2 1 , σ ξ = λ 1 . For the exponential distribution, the mean and standard deviation are therefore equal.
Moments of the Poisson distribution ¶ Let
ξ ∼ Pois ( a ) . \xi\sim\operatorname{Pois}(a). ξ ∼ Pois ( a ) . The probability mass function is
P ξ ( k ) = a k e − a k ! , k = 0 , 1 , 2 , … . P_\xi(k)
=
\frac{a^k e^{-a}}{k!},
\qquad
k=0,1,2,\ldots. P ξ ( k ) = k ! a k e − a , k = 0 , 1 , 2 , … . The expectation is
E ξ = ∑ k = 0 ∞ k a k e − a k ! = a e − a ∑ k = 1 ∞ a k − 1 ( k − 1 ) ! = a e − a e a = a . \begin{aligned}
\mathbb{E}\xi
&=
\sum_{k=0}^\infty
k\frac{a^k e^{-a}}{k!}\\
&=
ae^{-a}
\sum_{k=1}^\infty
\frac{a^{k-1}}{(k-1)!}\\
&=
ae^{-a}e^a\\
&=a.
\end{aligned} E ξ = k = 0 ∑ ∞ k k ! a k e − a = a e − a k = 1 ∑ ∞ ( k − 1 )! a k − 1 = a e − a e a = a . Thus,
E ξ = a . \boxed{\mathbb{E}\xi=a.} E ξ = a . A similar calculation gives
E ξ 2 = ∑ k = 0 ∞ k 2 a k e − a k ! = ∑ k = 0 ∞ [ k ( k − 1 ) + k ] a k e − a k ! = ∑ k = 0 ∞ k ( k − 1 ) a k e − a k ! + ∑ k = 0 ∞ k a k e − a k ! = ∑ k = 2 ∞ a k e − a ( k − 2 ) ! + a = a 2 e − a ∑ k = 2 ∞ a k − 2 ( k − 2 ) ! + a = a 2 e − a e a + a = a 2 + a . \begin{split}
\mathbb{E}\xi^2 = & \sum_{k=0}^\infty k^2 \frac{a^k e^{-a}}{k!} \\
= & \sum_{k=0}^\infty [k(k-1) + k] \frac{a^k e^{-a}}{k!} \\
= & \sum_{k=0}^\infty k(k-1) \frac{a^k e^{-a}}{k!} + \sum_{k=0}^\infty k \frac{a^k e^{-a}}{k!} \\
= & \sum_{k=2}^\infty \frac{a^k e^{-a}}{(k-2)!} + a \\
= & a^2 e^{-a} \sum_{k=2}^\infty \frac{a^{k-2}}{(k-2)!} + a \\
= & a^2 e^{-a} e^a + a \\
= & a^2 + a.
\end{split} E ξ 2 = = = = = = = k = 0 ∑ ∞ k 2 k ! a k e − a k = 0 ∑ ∞ [ k ( k − 1 ) + k ] k ! a k e − a k = 0 ∑ ∞ k ( k − 1 ) k ! a k e − a + k = 0 ∑ ∞ k k ! a k e − a k = 2 ∑ ∞ ( k − 2 )! a k e − a + a a 2 e − a k = 2 ∑ ∞ ( k − 2 )! a k − 2 + a a 2 e − a e a + a a 2 + a . Therefore,
Var ( ξ ) = E ξ 2 − ( E ξ ) 2 = a 2 + a − a 2 = a . \begin{aligned}
\operatorname{Var}(\xi)
&=
\mathbb{E}\xi^2-(\mathbb{E}\xi)^2\\
&=
a^2+a-a^2\\
&=a.
\end{aligned} Var ( ξ ) = E ξ 2 − ( E ξ ) 2 = a 2 + a − a 2 = a . Hence,
Var ( ξ ) = a . \boxed{
\operatorname{Var}(\xi)=a.
} Var ( ξ ) = a . Thus, for a Poisson random variable, the parameter a a a is simultaneously its expected value and its variance.
Chebyshev’s inequality ¶ Expectation and variance can also be used to obtain bounds on probabilities.
The result follows from the fact that, on the event
∣ ξ − E ξ ∣ ≥ ε , |\xi-\mathbb{E}\xi|\geq\varepsilon, ∣ ξ − E ξ ∣ ≥ ε , we have
( ξ − E ξ ) 2 ≥ ε 2 . (\xi-\mathbb{E}\xi)^2\geq\varepsilon^2. ( ξ − E ξ ) 2 ≥ ε 2 . Consequently,
ε 2 P ( { ∣ ξ − E ξ ∣ ≥ ε } ) ≤ E ( ξ − E ξ ) 2 = Var ( ξ ) . \varepsilon^2
\mathbb{P}
\left(
\{|\xi-\mathbb{E}\xi|\geq\varepsilon\}
\right)
\leq
\mathbb{E}
(\xi-\mathbb{E}\xi)^2
=
\operatorname{Var}(\xi). ε 2 P ( { ∣ ξ − E ξ ∣ ≥ ε } ) ≤ E ( ξ − E ξ ) 2 = Var ( ξ ) . Dividing by ε 2 \varepsilon^2 ε 2 gives the result.
Chebyshev’s inequality is generally not a sharp estimate, but it requires only the first two moments and makes no assumption about the particular distribution of ξ \xi ξ .
A simple application of Chebyshev’s inequality ¶ Suppose that a measurement ξ \xi ξ has
E ξ = 10 , Var ( ξ ) = 0.25. \mathbb{E}\xi=10,
\qquad
\operatorname{Var}(\xi)=0.25. E ξ = 10 , Var ( ξ ) = 0.25. We can bound the probability that the measurement differs from its expected value by at least 1:
P ( { ∣ ξ − 10 ∣ ≥ 1 } ) ≤ 0.25 1 2 = 0.25. \begin{aligned}
\mathbb{P}(\{|\xi-10|\geq1\})
&\leq
\frac{0.25}{1^2}\\
&=0.25.
\end{aligned} P ({ ∣ ξ − 10∣ ≥ 1 }) ≤ 1 2 0.25 = 0.25. Therefore,
P ( { 9 < ξ < 11 } ) ≥ 0.75. \mathbb{P}(\{9<\xi<11\})
\geq0.75. P ({ 9 < ξ < 11 }) ≥ 0.75. Notice that no assumption about the distribution of ξ \xi ξ was needed.
Generating functions ¶ For discrete random variables taking values in the nonnegative integers, an especially useful tool is the probability generating function .
Let
P ξ ( k ) = P ( { ξ = k } ) , k = 0 , 1 , 2 , … . P_\xi(k)
=
\mathbb{P}(\{\xi=k\}),
\qquad
k=0,1,2,\ldots. P ξ ( k ) = P ({ ξ = k }) , k = 0 , 1 , 2 , … . The probability generating function of ξ \xi ξ is defined by
F ξ ( z ) = ∑ k = 0 ∞ P ξ ( k ) z k , ∣ z ∣ ≤ 1. F_\xi(z)
=
\sum_{k=0}^{\infty}
P_\xi(k)z^k,
\qquad
|z|\leq1. F ξ ( z ) = k = 0 ∑ ∞ P ξ ( k ) z k , ∣ z ∣ ≤ 1. Because the probabilities sum to one,
F ξ ( 1 ) = ∑ k = 0 ∞ P ξ ( k ) = 1. F_\xi(1)
=
\sum_{k=0}^{\infty}P_\xi(k)
=
1. F ξ ( 1 ) = k = 0 ∑ ∞ P ξ ( k ) = 1. The generating function has a simple probabilistic interpretation:
F ξ ( z ) = E [ z ξ ] . F_\xi(z)
=
\mathbb{E}[z^\xi]. F ξ ( z ) = E [ z ξ ] . For ∣ z ∣ < 1 |z|<1 ∣ z ∣ < 1 , it can be differentiated term by term:
F ξ ′ ( z ) = ∑ k = 1 ∞ k P ξ ( k ) z k − 1 . F_\xi'(z)
=
\sum_{k=1}^{\infty}
kP_\xi(k)z^{k-1}. F ξ ′ ( z ) = k = 1 ∑ ∞ k P ξ ( k ) z k − 1 . Evaluating at z = 1 z=1 z = 1 gives
F ξ ′ ( 1 ) = E ξ , F_\xi'(1)
=
\mathbb{E}\xi, F ξ ′ ( 1 ) = E ξ , provided the first moment exists.
Similarly,
F ξ ′ ′ ( 1 ) = E [ ξ ( ξ − 1 ) ] . F_\xi''(1)
=
\mathbb{E}[\xi(\xi-1)]. F ξ ′′ ( 1 ) = E [ ξ ( ξ − 1 )] . Since
ξ 2 = ξ ( ξ − 1 ) + ξ , \xi^2=\xi(\xi-1)+\xi, ξ 2 = ξ ( ξ − 1 ) + ξ , we obtain
E ξ 2 = F ξ ′ ′ ( 1 ) + F ξ ′ ( 1 ) . \mathbb{E}\xi^2
=
F_\xi''(1)+F_\xi'(1). E ξ 2 = F ξ ′′ ( 1 ) + F ξ ′ ( 1 ) . Consequently,
Var ( ξ ) = F ξ ′ ′ ( 1 ) + F ξ ′ ( 1 ) − [ F ξ ′ ( 1 ) ] 2 . \boxed{
\operatorname{Var}(\xi)
=
F_\xi''(1)
+
F_\xi'(1)
-
[F_\xi'(1)]^2.
} Var ( ξ ) = F ξ ′′ ( 1 ) + F ξ ′ ( 1 ) − [ F ξ ′ ( 1 ) ] 2 . Generating functions therefore encode the moments of a discrete random variable.
Generating functions of important distributions ¶ The generating function of a Bernoulli random variable is particularly simple.
If
ξ ∼ Bern ( p ) , \xi\sim\operatorname{Bern}(p), ξ ∼ Bern ( p ) , then
F ξ ( z ) = ( 1 − p ) + p z . \begin{aligned}
F_\xi(z)
&=
(1-p)+pz.
\end{aligned} F ξ ( z ) = ( 1 − p ) + p z . Thus,
F ξ ( z ) = 1 − p + p z . F_\xi(z)=1-p+pz. F ξ ( z ) = 1 − p + p z . For a binomial random variable,
S n = ξ 1 + ⋯ + ξ n , S_n=\xi_1+\cdots+\xi_n, S n = ξ 1 + ⋯ + ξ n , where the ξ k \xi_k ξ k are independent Bernoulli variables, we have
F S n ( z ) = E [ z ξ 1 + ⋯ + ξ n ] = E [ z ξ 1 ⋯ z ξ n ] . \begin{aligned}
F_{S_n}(z)
&=
\mathbb{E}[z^{\xi_1+\cdots+\xi_n}]\\
&=
\mathbb{E}[z^{\xi_1}\cdots z^{\xi_n}].
\end{aligned} F S n ( z ) = E [ z ξ 1 + ⋯ + ξ n ] = E [ z ξ 1 ⋯ z ξ n ] . Independence gives
F S n ( z ) = ∏ k = 1 n E [ z ξ k ] = ( 1 − p + p z ) n . F_{S_n}(z)
=
\prod_{k=1}^n
\mathbb{E}[z^{\xi_k}]
=
(1-p+pz)^n. F S n ( z ) = k = 1 ∏ n E [ z ξ k ] = ( 1 − p + p z ) n . Therefore,
F S n ( z ) = ( 1 − p + p z ) n . \boxed{
F_{S_n}(z)
=
(1-p+pz)^n.
} F S n ( z ) = ( 1 − p + p z ) n . This is another way to see why the binomial distribution is associated with the binomial expansion.
The Poisson generating function ¶ For a Poisson random variable,
ξ ∼ Pois ( a ) , \xi\sim\operatorname{Pois}(a), ξ ∼ Pois ( a ) , we have
F ξ ( z ) = ∑ k = 0 ∞ a k e − a k ! z k = e − a ∑ k = 0 ∞ ( a z ) k k ! = e − a e a z . \begin{aligned}
F_\xi(z)
&=
\sum_{k=0}^{\infty}
\frac{a^k e^{-a}}{k!}z^k\\
&=
e^{-a}
\sum_{k=0}^{\infty}
\frac{(az)^k}{k!}\\
&=
e^{-a}e^{az}.
\end{aligned} F ξ ( z ) = k = 0 ∑ ∞ k ! a k e − a z k = e − a k = 0 ∑ ∞ k ! ( a z ) k = e − a e a z . Hence,
F ξ ( z ) = e a ( z − 1 ) . \boxed{
F_\xi(z)=e^{a(z-1)}.
} F ξ ( z ) = e a ( z − 1 ) . Differentiating,
F ξ ′ ( z ) = a e a ( z − 1 ) , F_\xi'(z)=ae^{a(z-1)}, F ξ ′ ( z ) = a e a ( z − 1 ) , and therefore
F ξ ′ ( 1 ) = a . F_\xi'(1)=a. F ξ ′ ( 1 ) = a . Similarly,
F ξ ′ ′ ( 1 ) = a 2 . F_\xi''(1)=a^2. F ξ ′′ ( 1 ) = a 2 . Using (40) ,
Var ( ξ ) = a 2 + a − a 2 = a . \operatorname{Var}(\xi)
=
a^2+a-a^2
=
a. Var ( ξ ) = a 2 + a − a 2 = a . The generating function therefore provides a particularly short derivation of the Poisson moments.
Why moments are useful ¶ The probability distribution of a random variable contains complete information about its probabilistic behavior.
In applications, however, it is often more convenient to work with a few summary quantities:
E ξ , Var ( ξ ) , E ( ξ − E ξ ) 3 , E ( ξ − E ξ ) 4 , … \mathbb{E}\xi,
\qquad
\operatorname{Var}(\xi),
\qquad
\mathbb{E}(\xi-\mathbb{E}\xi)^3,
\qquad
\mathbb{E}(\xi-\mathbb{E}\xi)^4,
\ldots E ξ , Var ( ξ ) , E ( ξ − E ξ ) 3 , E ( ξ − E ξ ) 4 , … The first moment describes location, while the second centered moment describes spread.
Higher moments can provide information about the shape of the distribution. For example:
We will not develop these quantities further here. The main point is that moments provide a systematic way of extracting numerical information from a probability distribution.
Averages of random variables ¶ One of the most important applications of expectation and variance concerns averages.
Let
ξ 1 , … , ξ n \xi_1,\ldots,\xi_n ξ 1 , … , ξ n be independent random variables with common expectation
E ξ k = a \mathbb{E}\xi_k=a E ξ k = a and common variance
Var ( ξ k ) = σ 2 . \operatorname{Var}(\xi_k)=\sigma^2. Var ( ξ k ) = σ 2 . Define their sample average by
ξ ‾ n = 1 n ∑ k = 1 n ξ k . \overline{\xi}_n
=
\frac1n\sum_{k=1}^n\xi_k. ξ n = n 1 k = 1 ∑ n ξ k . By linearity of expectation,
E ξ ‾ n = 1 n ∑ k = 1 n E ξ k = a . \mathbb{E}\overline{\xi}_n
=
\frac1n
\sum_{k=1}^n
\mathbb{E}\xi_k
=
a. E ξ n = n 1 k = 1 ∑ n E ξ k = a . Thus,
E ξ ‾ n = a . \boxed{
\mathbb{E}\overline{\xi}_n=a.
} E ξ n = a . For the variance,
Var ( ξ ‾ n ) = Var ( 1 n ∑ k = 1 n ξ k ) = 1 n 2 ∑ k = 1 n Var ( ξ k ) = σ 2 n . \begin{aligned}
\operatorname{Var}(\overline{\xi}_n)
&=
\operatorname{Var}
\left(
\frac1n\sum_{k=1}^n\xi_k
\right)\\
&=
\frac1{n^2}
\sum_{k=1}^n
\operatorname{Var}(\xi_k)\\
&=
\frac{\sigma^2}{n}.
\end{aligned} Var ( ξ n ) = Var ( n 1 k = 1 ∑ n ξ k ) = n 2 1 k = 1 ∑ n Var ( ξ k ) = n σ 2 . Therefore,
Var ( ξ ‾ n ) = σ 2 n . \boxed{
\operatorname{Var}(\overline{\xi}_n)
=
\frac{\sigma^2}{n}.
} Var ( ξ n ) = n σ 2 . This is an important result: averaging independent observations does not change their expected value, but it reduces their variance by a factor of n n n .
Equivalently,
sd ( ξ ‾ n ) = σ n . \operatorname{sd}(\overline{\xi}_n)
=
\frac{\sigma}{\sqrt{n}}. sd ( ξ n ) = n σ . This explains why averaging repeated measurements can substantially reduce random fluctuations.
Suppose that a quantity is measured repeatedly according to
X k = a + ε k , X_k=a+\varepsilon_k, X k = a + ε k , where the measurement errors ε k \varepsilon_k ε k are independent and satisfy
E ε k = 0 , Var ( ε k ) = σ 2 . \mathbb{E}\varepsilon_k=0,
\qquad
\operatorname{Var}(\varepsilon_k)=\sigma^2. E ε k = 0 , Var ( ε k ) = σ 2 . The average of n n n measurements is
X ‾ n = 1 n ∑ k = 1 n X k . \overline{X}_n
=
\frac1n\sum_{k=1}^n X_k. X n = n 1 k = 1 ∑ n X k . Since
E X k = a , \mathbb{E}X_k=a, E X k = a , we obtain
E X ‾ n = a . \mathbb{E}\overline{X}_n=a. E X n = a . Moreover,
Var ( X ‾ n ) = σ 2 n . \operatorname{Var}(\overline{X}_n)
=
\frac{\sigma^2}{n}. Var ( X n ) = n σ 2 . Thus, increasing the number of independent measurements reduces the variance of the average.
This distinction will become fundamental when we study the law of large numbers and the central limit theorem .
Summary ¶ In this lecture we introduced the main numerical characteristics of a random variable.
For a discrete random variable,
E ξ = ∑ x x P ξ ( x ) , \mathbb{E}\xi
=
\sum_xxP_\xi(x), E ξ = x ∑ x P ξ ( x ) , while for a continuous random variable,
E ξ = ∫ − ∞ + ∞ x p ξ ( x ) d x . \mathbb{E}\xi
=
\int_{-\infty}^{+\infty}xp_\xi(x)\,dx. E ξ = ∫ − ∞ + ∞ x p ξ ( x ) d x . More generally, expectation allows us to compute
E φ ( ξ ) . \mathbb{E}\varphi(\xi). E φ ( ξ ) . The main properties are linearity, monotonicity, and
∣ E ξ ∣ ≤ E ∣ ξ ∣ . |\mathbb{E}\xi|\leq\mathbb{E}|\xi|. ∣ E ξ ∣ ≤ E ∣ ξ ∣. The variance is
Var ( ξ ) = E [ ( ξ − E ξ ) 2 ] = E ξ 2 − ( E ξ ) 2 . \operatorname{Var}(\xi)
=
\mathbb{E}
\left[
(\xi-\mathbb{E}\xi)^2
\right]
=
\mathbb{E}\xi^2-(\mathbb{E}\xi)^2. Var ( ξ ) = E [ ( ξ − E ξ ) 2 ] = E ξ 2 − ( E ξ ) 2 . The standard deviation is
σ ξ = Var ( ξ ) . \sigma_\xi
=
\sqrt{\operatorname{Var}(\xi)}. σ ξ = Var ( ξ ) . We derived, among others,
ξ ∼ Bern ( p ) ⟹ E ξ = p , Var ( ξ ) = p ( 1 − p ) , ξ ∼ Bin ( n , p ) ⟹ E ξ = n p , Var ( ξ ) = n p ( 1 − p ) , ξ ∼ Pois ( a ) ⟹ E ξ = a , Var ( ξ ) = a , ξ ∼ Exp ( λ ) ⟹ E ξ = 1 λ , Var ( ξ ) = 1 λ 2 , ξ ∼ N ( a , σ 2 ) ⟹ E ξ = a , Var ( ξ ) = σ 2 . \begin{aligned}
\xi\sim\operatorname{Bern}(p)
&\quad\Longrightarrow\quad
\mathbb{E}\xi=p,
\qquad
\operatorname{Var}(\xi)=p(1-p),\\[4pt]
\xi\sim\operatorname{Bin}(n,p)
&\quad\Longrightarrow\quad
\mathbb{E}\xi=np,
\qquad
\operatorname{Var}(\xi)=np(1-p),\\[4pt]
\xi\sim\operatorname{Pois}(a)
&\quad\Longrightarrow\quad
\mathbb{E}\xi=a,
\qquad
\operatorname{Var}(\xi)=a,\\[4pt]
\xi\sim\operatorname{Exp}(\lambda)
&\quad\Longrightarrow\quad
\mathbb{E}\xi=\frac1\lambda,
\qquad
\operatorname{Var}(\xi)=\frac1{\lambda^2},\\[4pt]
\xi\sim N(a,\sigma^2)
&\quad\Longrightarrow\quad
\mathbb{E}\xi=a,
\qquad
\operatorname{Var}(\xi)=\sigma^2.
\end{aligned} ξ ∼ Bern ( p ) ξ ∼ Bin ( n , p ) ξ ∼ Pois ( a ) ξ ∼ Exp ( λ ) ξ ∼ N ( a , σ 2 ) ⟹ E ξ = p , Var ( ξ ) = p ( 1 − p ) , ⟹ E ξ = n p , Var ( ξ ) = n p ( 1 − p ) , ⟹ E ξ = a , Var ( ξ ) = a , ⟹ E ξ = λ 1 , Var ( ξ ) = λ 2 1 , ⟹ E ξ = a , Var ( ξ ) = σ 2 . We also introduced Chebyshev’s inequality,
P ( { ∣ ξ − E ξ ∣ ≥ ε } ) ≤ Var ( ξ ) ε 2 , \mathbb{P}
\left(
\{|\xi-\mathbb{E}\xi|\geq\varepsilon\}
\right)
\leq
\frac{\operatorname{Var}(\xi)}{\varepsilon^2}, P ( { ∣ ξ − E ξ ∣ ≥ ε } ) ≤ ε 2 Var ( ξ ) , and probability generating functions for nonnegative integer-valued random variables.
Finally, we saw that for the average of independent identically distributed random variables,
E ξ ‾ n = a , Var ( ξ ‾ n ) = σ 2 n . \mathbb{E}\overline{\xi}_n=a,
\qquad
\operatorname{Var}(\overline{\xi}_n)=\frac{\sigma^2}{n}. E ξ n = a , Var ( ξ n ) = n σ 2 . This reduction of fluctuations under averaging is one of the central ideas of probability theory.
In the next lecture we will extend these ideas to two or more random variables , introducing joint distributions, marginal and conditional distributions, covariance, correlation, and joint moments.