Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Lecture 6 - Expectation and Moments of a Random Variable

In the previous lectures we introduced random variables and several important probability distributions. We have seen how a probability distribution describes the possible values of a random variable and assigns probabilities to them.

It is often useful, however, to summarize a probability distribution by a small number of numerical quantities.

For example:

The most important quantities used for this purpose are the expectation, the moments, and the variance.

In this lecture we develop these concepts for a single random variable. We also introduce generating functions for discrete random variables, which provide a convenient way of computing moments and studying sums of independent random variables.


The expected value of a discrete random variable

Let ξ\xi be a discrete random variable with probability mass function PξP_\xi.

The expected value of ξ\xi is defined by

Eξ=∑xxPξ(x),\mathbb{E}\xi = \sum_x xP_\xi(x),

provided that the series is absolutely convergent.

The expectation can be interpreted as the long-run average value obtained by repeatedly observing the random experiment.

If the possible values of ξ\xi are

x1,x2,…x_1,x_2,\ldots

with probabilities

p1,p2,…,p_1,p_2,\ldots,

then

Eξ=x1p1+x2p2+⋯ .\mathbb{E}\xi = x_1p_1+x_2p_2+\cdots.

Thus, the expectation is a probability-weighted average of the possible values.

Expectation of a function of a random variable

Suppose that

η=φ(ξ)\eta=\varphi(\xi)

for some function φ\varphi.

There is no need to determine the probability distribution of η\eta before computing its expectation. Instead,

Eφ(ξ)=∑xφ(x)Pξ(x),\mathbb{E}\varphi(\xi) = \sum_x \varphi(x)P_\xi(x),

provided that the sum is well defined.

This result is sometimes called the law of the unconscious statistician.

For example, if ξ\xi is the outcome of a fair die and we want the expected value of ξ2\xi^2, then

Eξ2=∑k=16k216=1+4+9+16+25+366=916.\begin{aligned} \mathbb{E}\xi^2 &= \sum_{k=1}^6 k^2\frac16\\ &= \frac{1+4+9+16+25+36}{6}\\ &= \frac{91}{6}. \end{aligned}

We will see that this quantity is closely related to the variance of ξ\xi.


Expectation of a continuous random variable

For a continuous random variable with density pξp_\xi, the corresponding definition is obtained by replacing the sum with an integral:

Eξ=∫−∞+∞xpξ(x) dx,\mathbb{E}\xi = \int_{-\infty}^{+\infty} x p_\xi(x)\,dx,

provided that the integral is absolutely convergent.

More generally,

Eφ(ξ)=∫−∞+∞φ(x)pξ(x) dx.\mathbb{E}\varphi(\xi) = \int_{-\infty}^{+\infty} \varphi(x)p_\xi(x)\,dx.

Thus, the discrete and continuous cases have exactly the same structure:

Eφ(ξ)={∑xφ(x)Pξ(x),discrete,∫−∞+∞φ(x)pξ(x) dx,continuous.\boxed{ \mathbb{E}\varphi(\xi) = \begin{cases} \displaystyle \sum_x\varphi(x)P_\xi(x), & \text{discrete},\\[12pt] \displaystyle \int_{-\infty}^{+\infty} \varphi(x)p_\xi(x)\,dx, & \text{continuous}. \end{cases}}

Properties of expectation

Expectation satisfies several fundamental properties.

Constant

If cc is a constant, then

Ec=c.\mathbb{E}c=c.

Scaling

For any constant cc,

E(cξ)=cEξ.\mathbb{E}(c\xi) = c\mathbb{E}\xi.

Linearity

For random variables ξ\xi and η\eta for which the expectations exist,

E(ξ+η)=Eξ+Eη.\mathbb{E}(\xi+\eta) = \mathbb{E}\xi+\mathbb{E}\eta.

More generally,

E(c1ξ+c2η)=c1Eξ+c2Eη.\mathbb{E}(c_1\xi+c_2\eta) = c_1\mathbb{E}\xi+c_2\mathbb{E}\eta.

Importantly, independence is not required for linearity of expectation.

Monotonicity

If

ξ≤η\xi\leq\eta

with probability one, then

Eξ≤Eη.\mathbb{E}\xi\leq\mathbb{E}\eta.

Absolute value

The triangle inequality for expectation gives

∣Eξ∣≤E∣ξ∣.|\mathbb{E}\xi| \leq \mathbb{E}|\xi|.

This inequality is useful when establishing the existence of expectations and moments.


Expectation of a Bernoulli random variable

Let

ξ∼Bern⁡(p).\xi\sim\operatorname{Bern}(p).

Thus

Pξ(0)=1−p,Pξ(1)=p.P_\xi(0)=1-p, \qquad P_\xi(1)=p.

Its expectation is

Eξ=0(1−p)+1p=p.\begin{aligned} \mathbb{E}\xi &= 0(1-p)+1p\\ &=p. \end{aligned}

Therefore,

Eξ=p.\boxed{\mathbb{E}\xi=p.}

This simple result has an important interpretation: the expected value of an indicator variable is the probability of the event it represents.

If AA is an event, define its indicator by

1A={1,A occurs,0,A does not occur.\mathbf{1}_A = \begin{cases} 1, & A\text{ occurs},\\ 0, & A\text{ does not occur}. \end{cases}

Then

E1A=P(A).\mathbb{E}\mathbf{1}_A = \mathbb{P}(A).

Indicator variables will be particularly useful when studying sums of random variables.


Expectation of the binomial distribution

Recall that

Sn=ξ1+⋯+ξn,S_n=\xi_1+\cdots+\xi_n,

where the ξk\xi_k are independent Bernoulli random variables with parameter pp.

Since

Eξk=p,\mathbb{E}\xi_k=p,

linearity of expectation gives

ESn=E(ξ1+⋯+ξn)=∑k=1nEξk=np.\begin{aligned} \mathbb{E}S_n &= \mathbb{E}(\xi_1+\cdots+\xi_n)\\ &= \sum_{k=1}^n\mathbb{E}\xi_k\\ &= np. \end{aligned}

Hence, if

Sn∼Bin⁡(n,p),S_n\sim\operatorname{Bin}(n,p),

then

ESn=np.\boxed{\mathbb{E}S_n=np.}

The result does not require us to perform the binomial sum directly.

The interpretation is natural: if each of nn trials has success probability pp, then the expected number of successes is npnp.


Expectation of the Normal distribution

Let

ξ∼N(μ,σ2).\xi \sim \mathcal{N}(\mu, \sigma^2).

The probability density function is

fξ(x)=12πσ2e−(x−μ)22σ2,x∈R.f_\xi(x) = \frac{1}{\sqrt{2\pi\sigma^2}} e^{-\frac{(x-\mu)^2}{2\sigma^2}}, \qquad x \in \mathbb{R}.

The expectation is

Eξ=∫−∞∞xfξ(x) dx=∫−∞∞x12πσ2e−(x−μ)22σ2 dx=∫−∞∞(yσ+μ)12πσ2e−y22σ dy(substituting y=x−μσ, dx=σ dy)=σ∫−∞∞y12πe−y22 dy+μ∫−∞∞12πe−y22 dy=σ⋅0+μ⋅1=μ.\begin{split} \mathbb{E}\xi = & \int_{-\infty}^\infty x f_\xi(x) \, dx \\ = & \int_{-\infty}^\infty x \frac{1}{\sqrt{2\pi\sigma^2}} e^{-\frac{(x-\mu)^2}{2\sigma^2}} \, dx \\ = & \int_{-\infty}^\infty (y\sigma + \mu) \frac{1}{\sqrt{2\pi\sigma^2}} e^{-\frac{y^2}{2}} \sigma \, dy \quad \left(\text{substituting } y = \frac{x-\mu}{\sigma}, \, dx = \sigma \, dy\right) \\ = & \sigma \int_{-\infty}^\infty y \frac{1}{\sqrt{2\pi}} e^{-\frac{y^2}{2}} \, dy + \mu \int_{-\infty}^\infty \frac{1}{\sqrt{2\pi}} e^{-\frac{y^2}{2}} \, dy \\ = & \sigma \cdot 0 + \mu \cdot 1 \\ = & \mu. \end{split}

Thus,

Eξ=μ.\mathbb{E}\xi = \mu.

Moments

The kk-th raw moment of a random variable ξ\xi is the expectation

mk=Eξk,m_k = \mathbb{E}\xi^k,

whenever the expectation exists.

Thus,

m1=Eξ,m2=Eξ2,m3=Eξ3,m_1=\mathbb{E}\xi, \qquad m_2=\mathbb{E}\xi^2, \qquad m_3=\mathbb{E}\xi^3,

and so on.

The first moment is the expected value.

Higher moments provide additional information about the distribution.

For example, the second moment Eξ2\mathbb{E}\xi^2 is used to quantify the spread of a random variable, while higher moments can be used to characterize features such as asymmetry and tail behavior.

It is often useful to consider centered moments:

μk=E(ξ−Eξ)k.\mu_k = \mathbb{E}(\xi-\mathbb{E}\xi)^k.

The first centered moment is always zero:

E(ξ−Eξ)=0.\mathbb{E}(\xi-\mathbb{E}\xi)=0.

The second centered moment is the variance.


Variance

The variance of a random variable ξ\xi is defined by

Var⁡(ξ)=E[(ξ−Eξ)2].\operatorname{Var}(\xi) = \mathbb{E} \left[ (\xi-\mathbb{E}\xi)^2 \right].

It measures the average squared distance of ξ\xi from its expected value.

The variance is always nonnegative:

Var⁡(ξ)≥0.\operatorname{Var}(\xi)\geq0.

Indeed, the random variable

(ξ−Eξ)2(\xi-\mathbb{E}\xi)^2

is nonnegative.

A useful alternative expression follows by expanding the square:

Var⁡(ξ)=E[ξ2−2ξEξ+(Eξ)2]=Eξ2−2(Eξ)2+(Eξ)2.\begin{aligned} \operatorname{Var}(\xi) &= \mathbb{E} \left[ \xi^2-2\xi\mathbb{E}\xi+(\mathbb{E}\xi)^2 \right]\\ &= \mathbb{E}\xi^2 -2(\mathbb{E}\xi)^2 +(\mathbb{E}\xi)^2. \end{aligned}

Therefore,

Var⁡(ξ)=Eξ2−(Eξ)2.\boxed{ \operatorname{Var}(\xi) = \mathbb{E}\xi^2-(\mathbb{E}\xi)^2. }

This formula is often much more convenient for computations.


Variance of a Bernoulli random variable

Let

ξ∼Bern⁡(p).\xi\sim\operatorname{Bern}(p).

Since ξ\xi takes only the values 0 and 1,

ξ2=ξ.\xi^2=\xi.

Therefore,

Eξ2=Eξ=p.\mathbb{E}\xi^2=\mathbb{E}\xi=p.

Using (18),

Var⁡(ξ)=p−p2=p(1−p).\begin{aligned} \operatorname{Var}(\xi) &= p-p^2\\ &= p(1-p). \end{aligned}

Hence,

Var⁡(ξ)=p(1−p).\boxed{ \operatorname{Var}(\xi)=p(1-p). }

The variance is largest when p=1/2p=1/2 and becomes small when pp is close to either 0 or 1.


Variance of a sum of independent random variables

Let ξ\xi and η\eta be independent random variables with finite second moments.

We begin with

Var⁡(ξ+η)=E[(ξ+η−Eξ−Eη)2].\operatorname{Var}(\xi+\eta) = \mathbb{E} \left[ (\xi+\eta-\mathbb{E}\xi-\mathbb{E}\eta)^2 \right].

Expanding the square,

Var⁡(ξ+η)=E(ξ−Eξ)2+E(η−Eη)2+2E[(ξ−Eξ)(η−Eη)].\begin{aligned} \operatorname{Var}(\xi+\eta) &= \mathbb{E}(\xi-\mathbb{E}\xi)^2 + \mathbb{E}(\eta-\mathbb{E}\eta)^2\\ &\quad+ 2\mathbb{E} \left[ (\xi-\mathbb{E}\xi) (\eta-\mathbb{E}\eta) \right]. \end{aligned}

Independence implies

E[(ξ−Eξ)(η−Eη)]=E(ξ−Eξ)E(η−Eη)=0.\mathbb{E} \left[ (\xi-\mathbb{E}\xi) (\eta-\mathbb{E}\eta) \right] = \mathbb{E}(\xi-\mathbb{E}\xi) \mathbb{E}(\eta-\mathbb{E}\eta) = 0.

Consequently,

Var⁡(ξ+η)=Var⁡(ξ)+Var⁡(η).\boxed{ \operatorname{Var}(\xi+\eta) = \operatorname{Var}(\xi) + \operatorname{Var}(\eta). }

By induction, if ξ1,…,ξn\xi_1,\ldots,\xi_n are independent,

Var⁡(∑k=1nξk)=∑k=1nVar⁡(ξk).\boxed{ \operatorname{Var} \left( \sum_{k=1}^n\xi_k \right) = \sum_{k=1}^n\operatorname{Var}(\xi_k). }

Variance of the binomial distribution

For

Sn=ξ1+⋯+ξn,S_n=\xi_1+\cdots+\xi_n,

with independent Bernoulli random variables,

Var⁡(ξk)=p(1−p).\operatorname{Var}(\xi_k)=p(1-p).

Therefore, (21) gives

Var⁡(Sn)=∑k=1np(1−p)=np(1−p).\begin{aligned} \operatorname{Var}(S_n) &= \sum_{k=1}^n p(1-p)\\ &= np(1-p). \end{aligned}

Thus,

Var⁡(Sn)=np(1−p).\boxed{ \operatorname{Var}(S_n)=np(1-p). }

We have therefore obtained both parameters that appeared in the normal approximation of Lecture 5:

ESn=np,Var⁡(Sn)=np(1−p).\mathbb{E}S_n=np, \qquad \operatorname{Var}(S_n)=np(1-p).

Variance of the Normal distribution

Let ξ∼N(μ,σ2)\xi \sim \mathcal{N}(\mu, \sigma^2), as for the other cases it is useful to use the relation

Var⁡(ξ)=E(ξ−μ)2=Eξ2−(Eξ)2.\operatorname{Var}(\xi) = \mathbb{E}(\xi - \mu)^2 = \mathbb{E}\xi^2 - (\mathbb{E}\xi)^2.

where we have used the fact that Eξ=μE\xi = \mu for the Normal distribution.

We justt to compute the second order moment Eξ2\mathbb{E}\xi^2 as

Eξ2=∫−∞∞x212πσ2e−(x−μ)22σ2 dx=∫−∞∞(yσ+μ)212πe−y22 dy(substituting y=x−μσ, dx=σ dy)=∫−∞∞(y2σ2+2μσy+μ2)12πe−y22 dy=σ2∫−∞∞y212πe−y22 dy+2μσ∫−∞∞y12πe−y22 dy+μ2∫−∞∞12πe−y22 dy=σ2([−y12πe−y22]−∞∞−∫−∞∞(−12πe−y22)dy)+2μσ⋅0+μ2⋅1=σ2(0+∫−∞∞12πe−y22 dy)+μ2=σ2⋅1+μ2=μ2+σ2.\begin{split} \mathbb{E}\xi^2 = & \int_{-\infty}^\infty x^2 \frac{1}{\sqrt{2\pi\sigma^2}} e^{-\frac{(x-\mu)^2}{2\sigma^2}} \, dx \\ = & \int_{-\infty}^\infty (y\sigma + \mu)^2 \frac{1}{\sqrt{2\pi}} e^{-\frac{y^2}{2}} \, dy \quad \left(\text{substituting } y = \frac{x-\mu}{\sigma}, \, dx = \sigma \, dy\right) \\ = & \int_{-\infty}^\infty (y^2 \sigma^2 + 2\mu\sigma y + \mu^2) \frac{1}{\sqrt{2\pi}} e^{-\frac{y^2}{2}} \, dy \\ = & \sigma^2 \int_{-\infty}^\infty y^2 \frac{1}{\sqrt{2\pi}} e^{-\frac{y^2}{2}} \, dy + 2\mu\sigma \int_{-\infty}^\infty y \frac{1}{\sqrt{2\pi}} e^{-\frac{y^2}{2}} \, dy + \mu^2 \int_{-\infty}^\infty \frac{1}{\sqrt{2\pi}} e^{-\frac{y^2}{2}} \, dy \\ = & \sigma^2 \left( \left[ -y \frac{1}{\sqrt{2\pi}} e^{-\frac{y^2}{2}} \right]_{-\infty}^\infty - \int_{-\infty}^\infty \left( -\frac{1}{\sqrt{2\pi}} e^{-\frac{y^2}{2}} \right) dy \right) + 2\mu\sigma \cdot 0 + \mu^2 \cdot 1 \\ = & \sigma^2 \left( 0 + \int_{-\infty}^\infty \frac{1}{\sqrt{2\pi}} e^{-\frac{y^2}{2}} \, dy \right) + \mu^2 \\ = & \sigma^2 \cdot 1 + \mu^2 \\ = & \mu^2 + \sigma^2. \end{split}

Now, substituting Eξ2=μ2+σ2\mathbb{E}\xi^2 = \mu^2 + \sigma^2:

Var⁡(ξ)=Eξ2−(Eξ)2=(μ2+σ2)−μ2=σ2.\begin{split} \operatorname{Var}(\xi) = & \mathbb{E}\xi^2 - (\mathbb{E}\xi)^2 \\ = & (\mu^2 + \sigma^2) - \mu^2 \\ = & \sigma^2. \end{split}

Thus,

Var⁡(ξ)=σ2.\operatorname{Var}(\xi) = \sigma^2.

Standard deviation

The variance has the units of the square of the original quantity.

For example, if ξ\xi represents a length measured in meters, then Var⁡(ξ)\operatorname{Var}(\xi) is measured in square meters.

It is often more natural to measure spread in the same units as the random variable. This leads to the standard deviation:

σξ=Var⁡(ξ).\sigma_\xi = \sqrt{\operatorname{Var}(\xi)}.

Thus, the standard deviation is the positive square root of the variance.

The standard deviation measures the typical scale of fluctuations around the expected value.

For a binomial random variable,

σSn=np(1−p).\sigma_{S_n} = \sqrt{np(1-p)}.

This is precisely the scaling that appeared in the De Moivre--Laplace theorem.


Affine transformations

Suppose that

η=cξ+d,\eta=c\xi+d,

where cc and dd are constants.

Using linearity of expectation,

Eη=cEξ+d.\mathbb{E}\eta = c\mathbb{E}\xi+d.

For the variance,

Var⁡(η)=Var⁡(cξ+d)=E[cξ+d−(cEξ+d)]2=E[c(ξ−Eξ)]2=c2Var⁡(ξ).\begin{aligned} \operatorname{Var}(\eta) &= \operatorname{Var}(c\xi+d)\\ &= \mathbb{E} \left[ c\xi+d-(c\mathbb{E}\xi+d) \right]^2\\ &= \mathbb{E} \left[ c(\xi-\mathbb{E}\xi) \right]^2\\ &= c^2\operatorname{Var}(\xi). \end{aligned}

Therefore,

Var⁡(cξ+d)=c2Var⁡(ξ).\boxed{ \operatorname{Var}(c\xi+d) = c^2\operatorname{Var}(\xi). }

In particular, adding a constant changes the mean but does not change the variance, while multiplying by cc scales the standard deviation by ∣c∣|c|.


Moments of the normal distribution

Consider

ξ∼N(a,σ2).\xi\sim N(a,\sigma^2).

From Lecture 5, we can write

ξ=a+σZ,\xi=a+\sigma Z,

where

Z∼N(0,1).Z\sim N(0,1).

The standard normal distribution is symmetric about zero. Consequently,

EZ=0.\mathbb{E}Z=0.

Its variance is

Var⁡(Z)=1.\operatorname{Var}(Z)=1.

Therefore, using (25) and (26),

Eξ=a,Var⁡(ξ)=σ2.\boxed{ \mathbb{E}\xi=a, \qquad \operatorname{Var}(\xi)=\sigma^2. }

This gives the probabilistic interpretation of the two parameters of the normal distribution:


Moments of the exponential distribution

Consider

ξ∼Exp⁡(λ).\xi\sim\operatorname{Exp}(\lambda).

Its expectation is

Eξ=∫0∞xλe−λx dx.\begin{aligned} \mathbb{E}\xi &= \int_0^\infty x\lambda e^{-\lambda x}\,dx. \end{aligned}

Integration by parts gives

∫0∞xλe−λx dx=[−xe−λx]0∞−∫0∞(−e−λx)dx=lim⁡x→∞−xeλx+0⋅e−λ⋅0+[−1λe−λx]0∞=lim⁡x→∞(−1λe−λx)+1λe−λ0=1λ.\begin{split} \int_0^\infty x \lambda e^{-\lambda x} \, dx = & \left[ -x e^{-\lambda x} \right]_0^\infty - \int_0^\infty \left( -e^{-\lambda x} \right) dx \\ = & \lim_{x \to \infty} \frac{-x}{e^{\lambda x}} + 0 \cdot e^{-\lambda \cdot 0} + \left[ -\frac{1}{\lambda} e^{-\lambda x} \right]_0^\infty \\ = & \lim_{x \to \infty} \left( -\frac{1}{\lambda} e^{-\lambda x} \right) + \frac{1}{\lambda} e^{-\lambda 0} \\ = & \frac{1}{\lambda}. \end{split}

Hence,

Eξ=1λ.\boxed{ \mathbb{E}\xi=\frac1\lambda. }

The second moment is

Eξ2=∫0∞x2λe−λx dx=[−x2e−λx]0∞−∫0∞2x(−e−λx)dx=lim⁡x→∞−x2eλx+02⋅e−λ⋅0+2λ∫0∞xλe−λx dx=0+0+2λ([−xe−λx]0∞−∫0∞(−e−λx)dx)=2λ(lim⁡x→∞−xeλx+0⋅e−λ⋅0+[−1λe−λx]0∞)=2λ(lim⁡x→∞(−1λe−λx)+1λe−λ⋅0)=2λ2.\begin{split} \mathbb{E}\xi^2 = & \int_0^\infty x^2 \lambda e^{-\lambda x} \, dx \\ = & \left[ -x^2 e^{-\lambda x} \right]_0^\infty - \int_0^\infty 2x \left( -e^{-\lambda x} \right) dx \\ = & \lim_{x \to \infty} \frac{-x^2}{e^{\lambda x}} + 0^2 \cdot e^{-\lambda \cdot 0} + \frac{2}{\lambda} \int_0^\infty x \lambda e^{-\lambda x} \, dx \\ = & 0 + 0 + \frac{2}{\lambda} \left( \left[ -x e^{-\lambda x} \right]_0^\infty - \int_0^\infty \left( -e^{-\lambda x} \right) dx \right) \\ = & \frac{2}{\lambda} \left( \lim_{x \to \infty} \frac{-x}{e^{\lambda x}} + 0 \cdot e^{-\lambda \cdot 0} + \left[ -\frac{1}{\lambda} e^{-\lambda x} \right]_0^\infty \right) \\ = & \frac{2}{\lambda} \left( \lim_{x \to \infty} \left( -\frac{1}{\lambda} e^{-\lambda x} \right) + \frac{1}{\lambda} e^{-\lambda \cdot 0} \right) \\ = & \frac{2}{\lambda^2}. \end{split}

Therefore,

Var⁡(ξ)=Eξ2−(Eξ)2=2λ2−1λ2=1λ2.\begin{aligned} \operatorname{Var}(\xi) &= \mathbb{E}\xi^2-(\mathbb{E}\xi)^2\\ &= \frac{2}{\lambda^2} - \frac{1}{\lambda^2}\\ &= \frac{1}{\lambda^2}. \end{aligned}

Thus,

Var⁡(ξ)=1λ2,σξ=1λ.\boxed{ \operatorname{Var}(\xi)=\frac1{\lambda^2}, \qquad \sigma_\xi=\frac1\lambda. }

For the exponential distribution, the mean and standard deviation are therefore equal.


Moments of the Poisson distribution

Let

ξ∼Pois⁡(a).\xi\sim\operatorname{Pois}(a).

The probability mass function is

Pξ(k)=ake−ak!,k=0,1,2,….P_\xi(k) = \frac{a^k e^{-a}}{k!}, \qquad k=0,1,2,\ldots.

The expectation is

Eξ=∑k=0∞kake−ak!=ae−a∑k=1∞ak−1(k−1)!=ae−aea=a.\begin{aligned} \mathbb{E}\xi &= \sum_{k=0}^\infty k\frac{a^k e^{-a}}{k!}\\ &= ae^{-a} \sum_{k=1}^\infty \frac{a^{k-1}}{(k-1)!}\\ &= ae^{-a}e^a\\ &=a. \end{aligned}

Thus,

Eξ=a.\boxed{\mathbb{E}\xi=a.}

A similar calculation gives

Eξ2=∑k=0∞k2ake−ak!=∑k=0∞[k(k−1)+k]ake−ak!=∑k=0∞k(k−1)ake−ak!+∑k=0∞kake−ak!=∑k=2∞ake−a(k−2)!+a=a2e−a∑k=2∞ak−2(k−2)!+a=a2e−aea+a=a2+a.\begin{split} \mathbb{E}\xi^2 = & \sum_{k=0}^\infty k^2 \frac{a^k e^{-a}}{k!} \\ = & \sum_{k=0}^\infty [k(k-1) + k] \frac{a^k e^{-a}}{k!} \\ = & \sum_{k=0}^\infty k(k-1) \frac{a^k e^{-a}}{k!} + \sum_{k=0}^\infty k \frac{a^k e^{-a}}{k!} \\ = & \sum_{k=2}^\infty \frac{a^k e^{-a}}{(k-2)!} + a \\ = & a^2 e^{-a} \sum_{k=2}^\infty \frac{a^{k-2}}{(k-2)!} + a \\ = & a^2 e^{-a} e^a + a \\ = & a^2 + a. \end{split}

Therefore,

Var⁡(ξ)=Eξ2−(Eξ)2=a2+a−a2=a.\begin{aligned} \operatorname{Var}(\xi) &= \mathbb{E}\xi^2-(\mathbb{E}\xi)^2\\ &= a^2+a-a^2\\ &=a. \end{aligned}

Hence,

Var⁡(ξ)=a.\boxed{ \operatorname{Var}(\xi)=a. }

Thus, for a Poisson random variable, the parameter aa is simultaneously its expected value and its variance.


Chebyshev’s inequality

Expectation and variance can also be used to obtain bounds on probabilities.

The result follows from the fact that, on the event

∣ξ−Eξ∣≥ε,|\xi-\mathbb{E}\xi|\geq\varepsilon,

we have

(ξ−Eξ)2≥ε2.(\xi-\mathbb{E}\xi)^2\geq\varepsilon^2.

Consequently,

ε2P({∣ξ−Eξ∣≥ε})≤E(ξ−Eξ)2=Var⁡(ξ).\varepsilon^2 \mathbb{P} \left( \{|\xi-\mathbb{E}\xi|\geq\varepsilon\} \right) \leq \mathbb{E} (\xi-\mathbb{E}\xi)^2 = \operatorname{Var}(\xi).

Dividing by ε2\varepsilon^2 gives the result.

Chebyshev’s inequality is generally not a sharp estimate, but it requires only the first two moments and makes no assumption about the particular distribution of ξ\xi.


A simple application of Chebyshev’s inequality

Suppose that a measurement ξ\xi has

Eξ=10,Var⁡(ξ)=0.25.\mathbb{E}\xi=10, \qquad \operatorname{Var}(\xi)=0.25.

We can bound the probability that the measurement differs from its expected value by at least 1:

P({∣ξ−10∣≥1})≤0.2512=0.25.\begin{aligned} \mathbb{P}(\{|\xi-10|\geq1\}) &\leq \frac{0.25}{1^2}\\ &=0.25. \end{aligned}

Therefore,

P({9<ξ<11})≥0.75.\mathbb{P}(\{9<\xi<11\}) \geq0.75.

Notice that no assumption about the distribution of ξ\xi was needed.


Generating functions

For discrete random variables taking values in the nonnegative integers, an especially useful tool is the probability generating function.

Let

Pξ(k)=P({ξ=k}),k=0,1,2,….P_\xi(k) = \mathbb{P}(\{\xi=k\}), \qquad k=0,1,2,\ldots.

The probability generating function of ξ\xi is defined by

Fξ(z)=∑k=0∞Pξ(k)zk,∣z∣≤1.F_\xi(z) = \sum_{k=0}^{\infty} P_\xi(k)z^k, \qquad |z|\leq1.

Because the probabilities sum to one,

Fξ(1)=∑k=0∞Pξ(k)=1.F_\xi(1) = \sum_{k=0}^{\infty}P_\xi(k) = 1.

The generating function has a simple probabilistic interpretation:

Fξ(z)=E[zξ].F_\xi(z) = \mathbb{E}[z^\xi].

For ∣z∣<1|z|<1, it can be differentiated term by term:

Fξ′(z)=∑k=1∞kPξ(k)zk−1.F_\xi'(z) = \sum_{k=1}^{\infty} kP_\xi(k)z^{k-1}.

Evaluating at z=1z=1 gives

Fξ′(1)=Eξ,F_\xi'(1) = \mathbb{E}\xi,

provided the first moment exists.

Similarly,

Fξ′′(1)=E[ξ(ξ−1)].F_\xi''(1) = \mathbb{E}[\xi(\xi-1)].

Since

ξ2=ξ(ξ−1)+ξ,\xi^2=\xi(\xi-1)+\xi,

we obtain

Eξ2=Fξ′′(1)+Fξ′(1).\mathbb{E}\xi^2 = F_\xi''(1)+F_\xi'(1).

Consequently,

Var⁡(ξ)=Fξ′′(1)+Fξ′(1)−[Fξ′(1)]2.\boxed{ \operatorname{Var}(\xi) = F_\xi''(1) + F_\xi'(1) - [F_\xi'(1)]^2. }

Generating functions therefore encode the moments of a discrete random variable.


Generating functions of important distributions

The generating function of a Bernoulli random variable is particularly simple.

If

ξ∼Bern⁡(p),\xi\sim\operatorname{Bern}(p),

then

Fξ(z)=(1−p)+pz.\begin{aligned} F_\xi(z) &= (1-p)+pz. \end{aligned}

Thus,

Fξ(z)=1−p+pz.F_\xi(z)=1-p+pz.

For a binomial random variable,

Sn=ξ1+⋯+ξn,S_n=\xi_1+\cdots+\xi_n,

where the ξk\xi_k are independent Bernoulli variables, we have

FSn(z)=E[zξ1+⋯+ξn]=E[zξ1⋯zξn].\begin{aligned} F_{S_n}(z) &= \mathbb{E}[z^{\xi_1+\cdots+\xi_n}]\\ &= \mathbb{E}[z^{\xi_1}\cdots z^{\xi_n}]. \end{aligned}

Independence gives

FSn(z)=∏k=1nE[zξk]=(1−p+pz)n.F_{S_n}(z) = \prod_{k=1}^n \mathbb{E}[z^{\xi_k}] = (1-p+pz)^n.

Therefore,

FSn(z)=(1−p+pz)n.\boxed{ F_{S_n}(z) = (1-p+pz)^n. }

This is another way to see why the binomial distribution is associated with the binomial expansion.


The Poisson generating function

For a Poisson random variable,

ξ∼Pois⁡(a),\xi\sim\operatorname{Pois}(a),

we have

Fξ(z)=∑k=0∞ake−ak!zk=e−a∑k=0∞(az)kk!=e−aeaz.\begin{aligned} F_\xi(z) &= \sum_{k=0}^{\infty} \frac{a^k e^{-a}}{k!}z^k\\ &= e^{-a} \sum_{k=0}^{\infty} \frac{(az)^k}{k!}\\ &= e^{-a}e^{az}. \end{aligned}

Hence,

Fξ(z)=ea(z−1).\boxed{ F_\xi(z)=e^{a(z-1)}. }

Differentiating,

Fξ′(z)=aea(z−1),F_\xi'(z)=ae^{a(z-1)},

and therefore

Fξ′(1)=a.F_\xi'(1)=a.

Similarly,

Fξ′′(1)=a2.F_\xi''(1)=a^2.

Using (40),

Var⁡(ξ)=a2+a−a2=a.\operatorname{Var}(\xi) = a^2+a-a^2 = a.

The generating function therefore provides a particularly short derivation of the Poisson moments.


Why moments are useful

The probability distribution of a random variable contains complete information about its probabilistic behavior.

In applications, however, it is often more convenient to work with a few summary quantities:

Eξ,Var⁡(ξ),E(ξ−Eξ)3,E(ξ−Eξ)4,…\mathbb{E}\xi, \qquad \operatorname{Var}(\xi), \qquad \mathbb{E}(\xi-\mathbb{E}\xi)^3, \qquad \mathbb{E}(\xi-\mathbb{E}\xi)^4, \ldots

The first moment describes location, while the second centered moment describes spread.

Higher moments can provide information about the shape of the distribution. For example:

We will not develop these quantities further here. The main point is that moments provide a systematic way of extracting numerical information from a probability distribution.


Averages of random variables

One of the most important applications of expectation and variance concerns averages.

Let

ξ1,…,ξn\xi_1,\ldots,\xi_n

be independent random variables with common expectation

Eξk=a\mathbb{E}\xi_k=a

and common variance

Var⁡(ξk)=σ2.\operatorname{Var}(\xi_k)=\sigma^2.

Define their sample average by

ξ‾n=1n∑k=1nξk.\overline{\xi}_n = \frac1n\sum_{k=1}^n\xi_k.

By linearity of expectation,

Eξ‾n=1n∑k=1nEξk=a.\mathbb{E}\overline{\xi}_n = \frac1n \sum_{k=1}^n \mathbb{E}\xi_k = a.

Thus,

Eξ‾n=a.\boxed{ \mathbb{E}\overline{\xi}_n=a. }

For the variance,

Var⁡(ξ‾n)=Var⁡(1n∑k=1nξk)=1n2∑k=1nVar⁡(ξk)=σ2n.\begin{aligned} \operatorname{Var}(\overline{\xi}_n) &= \operatorname{Var} \left( \frac1n\sum_{k=1}^n\xi_k \right)\\ &= \frac1{n^2} \sum_{k=1}^n \operatorname{Var}(\xi_k)\\ &= \frac{\sigma^2}{n}. \end{aligned}

Therefore,

Var⁡(ξ‾n)=σ2n.\boxed{ \operatorname{Var}(\overline{\xi}_n) = \frac{\sigma^2}{n}. }

This is an important result: averaging independent observations does not change their expected value, but it reduces their variance by a factor of nn.

Equivalently,

sd⁡(ξ‾n)=σn.\operatorname{sd}(\overline{\xi}_n) = \frac{\sigma}{\sqrt{n}}.

This explains why averaging repeated measurements can substantially reduce random fluctuations.

Summary

In this lecture we introduced the main numerical characteristics of a random variable.

For a discrete random variable,

Eξ=∑xxPξ(x),\mathbb{E}\xi = \sum_xxP_\xi(x),

while for a continuous random variable,

Eξ=∫−∞+∞xpξ(x) dx.\mathbb{E}\xi = \int_{-\infty}^{+\infty}xp_\xi(x)\,dx.

More generally, expectation allows us to compute

Eφ(ξ).\mathbb{E}\varphi(\xi).

The main properties are linearity, monotonicity, and

∣Eξ∣≤E∣ξ∣.|\mathbb{E}\xi|\leq\mathbb{E}|\xi|.

The variance is

Var⁡(ξ)=E[(ξ−Eξ)2]=Eξ2−(Eξ)2.\operatorname{Var}(\xi) = \mathbb{E} \left[ (\xi-\mathbb{E}\xi)^2 \right] = \mathbb{E}\xi^2-(\mathbb{E}\xi)^2.

The standard deviation is

σξ=Var⁡(ξ).\sigma_\xi = \sqrt{\operatorname{Var}(\xi)}.

We derived, among others,

ξ∼Bern⁡(p)⟹Eξ=p,Var⁡(ξ)=p(1−p),ξ∼Bin⁡(n,p)⟹Eξ=np,Var⁡(ξ)=np(1−p),ξ∼Pois⁡(a)⟹Eξ=a,Var⁡(ξ)=a,ξ∼Exp⁡(λ)⟹Eξ=1λ,Var⁡(ξ)=1λ2,ξ∼N(a,σ2)⟹Eξ=a,Var⁡(ξ)=σ2.\begin{aligned} \xi\sim\operatorname{Bern}(p) &\quad\Longrightarrow\quad \mathbb{E}\xi=p, \qquad \operatorname{Var}(\xi)=p(1-p),\\[4pt] \xi\sim\operatorname{Bin}(n,p) &\quad\Longrightarrow\quad \mathbb{E}\xi=np, \qquad \operatorname{Var}(\xi)=np(1-p),\\[4pt] \xi\sim\operatorname{Pois}(a) &\quad\Longrightarrow\quad \mathbb{E}\xi=a, \qquad \operatorname{Var}(\xi)=a,\\[4pt] \xi\sim\operatorname{Exp}(\lambda) &\quad\Longrightarrow\quad \mathbb{E}\xi=\frac1\lambda, \qquad \operatorname{Var}(\xi)=\frac1{\lambda^2},\\[4pt] \xi\sim N(a,\sigma^2) &\quad\Longrightarrow\quad \mathbb{E}\xi=a, \qquad \operatorname{Var}(\xi)=\sigma^2. \end{aligned}

We also introduced Chebyshev’s inequality,

P({∣ξ−Eξ∣≥ε})≤Var⁡(ξ)ε2,\mathbb{P} \left( \{|\xi-\mathbb{E}\xi|\geq\varepsilon\} \right) \leq \frac{\operatorname{Var}(\xi)}{\varepsilon^2},

and probability generating functions for nonnegative integer-valued random variables.

Finally, we saw that for the average of independent identically distributed random variables,

Eξ‾n=a,Var⁡(ξ‾n)=σ2n.\mathbb{E}\overline{\xi}_n=a, \qquad \operatorname{Var}(\overline{\xi}_n)=\frac{\sigma^2}{n}.

This reduction of fluctuations under averaging is one of the central ideas of probability theory.

In the next lecture we will extend these ideas to two or more random variables, introducing joint distributions, marginal and conditional distributions, covariance, correlation, and joint moments.