Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Lecture 8 - Law of Large Numbers and Central Limit Theorem

In the previous lectures we introduced random variables, their distributions, moments, and the joint characterization of pairs of random variables.

A natural next step is to consider a sequence of observations

ξ1,ξ2,…,ξn,…\xi_1,\xi_2,\ldots,\xi_n,\ldots

and to study quantities such as their sum

Sn=∑k=1nξkS_n=\sum_{k=1}^n \xi_k

and their average

ξ‾n=1n∑k=1nξk=Snn.\overline{\xi}_n = \frac{1}{n}\sum_{k=1}^n\xi_k = \frac{S_n}{n}.

These quantities arise naturally whenever an experiment is repeated.

For example:

Suppose that the random variables ξ1,…,ξn\xi_1,\ldots,\xi_n are independent and identically distributed, with

E[ξk]=μ,Var⁡(ξk)=σ2.\mathbb{E}[\xi_k]=\mu, \qquad \operatorname{Var}(\xi_k)=\sigma^2.

From the linearity of expectation,

E[ξ‾n]=μ.\mathbb{E}[\overline{\xi}_n] = \mu.

If the variables are independent, then

Var⁡(ξ‾n)=σ2n.\operatorname{Var}(\overline{\xi}_n) = \frac{\sigma^2}{n}.

Consequently,

SD⁡(ξ‾n)=σn.\operatorname{SD}(\overline{\xi}_n) = \frac{\sigma}{\sqrt{n}}.

The average therefore becomes increasingly concentrated around μ\mu as nn increases.

This observation leads to two fundamental questions:

  1. Does ξ‾n\overline{\xi}_n actually converge to μ\mu as n→∞n\to\infty?

  2. If it does, how large are the fluctuations of ξ‾n\overline{\xi}_n around μ\mu?

The first question is answered by the Law of Large Numbers.

The second is answered by the Central Limit Theorem.

Convergence in Probability

To formulate the Law of Large Numbers, we first need a notion of convergence for random variables.

Let ξ1,ξ2,…\xi_1,\xi_2,\ldots and ξ\xi be random variables. We say that ξn\xi_n converges in probability to ξ\xi if, for every ε>0\varepsilon>0,

lim⁡n→∞P({∣ξn−ξ∣>ε})=0.\lim_{n\to\infty} \mathbb{P} \left( \left\{ |\xi_n-\xi|>\varepsilon \right\} \right) = 0.

We write

ξn→Pξ.\xi_n\xrightarrow{\mathbb{P}}\xi.

The interpretation is that, for every fixed tolerance ε>0\varepsilon>0, the probability that ξn\xi_n differs from ξ\xi by more than ε\varepsilon becomes arbitrarily small.

For example, if

Var⁡(ξn)⟶0\operatorname{Var}(\xi_n)\longrightarrow 0

and

E[ξn]⟶μ,\mathbb{E}[\xi_n]\longrightarrow \mu,

then one expects ξn\xi_n to become concentrated around μ\mu. Chebyshev’s inequality makes this precise.

A Useful Consequence of Chebyshev’s Inequality

Recall Chebyshev’s inequality from Lecture 6:

P({∣ξ−E[ξ]∣≥ε})≤Var⁡(ξ)ε2.\mathbb{P} \left( \left\{ |\xi-\mathbb{E}[\xi]|\geq\varepsilon \right\} \right) \leq \frac{\operatorname{Var}(\xi)}{\varepsilon^2}.

Suppose that

E[ξn]=μ\mathbb{E}[\xi_n]=\mu

for every nn, and

Var⁡(ξn)⟶0.\operatorname{Var}(\xi_n)\longrightarrow 0.

Then

P(∣ξn−μ∣≥ε)≤Var⁡(ξn)ε2⟶0.\mathbb{P} \left( |\xi_n-\mu|\geq\varepsilon \right) \leq \frac{\operatorname{Var}(\xi_n)}{\varepsilon^2} \longrightarrow 0.

Therefore,

ξn→Pμ.\xi_n\xrightarrow{\mathbb{P}}\mu.

This simple argument is the main tool behind the elementary form of the Law of Large Numbers.

The Weak Law of Large Numbers

We can now state the fundamental result.

The theorem says that the average of a large number of independent observations approaches the common expected value.

This provides a mathematical foundation for the idea that repeated measurements can reveal an underlying mean.

The Law of Large Numbers for Bernoulli Trials

Consider independent Bernoulli random variables

ξk={1,success,0,failure,\xi_k= \begin{cases} 1,&\text{success},\\ 0,&\text{failure}, \end{cases}

with

P({ξk=1})=p.\mathbb{P}(\{\xi_k=1\})=p.

The sum

Sn=∑k=1nξkS_n=\sum_{k=1}^n\xi_k

counts the number of successes in nn trials.

The fraction of successes is

Snn=ξ‾n.\frac{S_n}{n} = \overline{\xi}_n.

Since

E[ξk]=p,\mathbb{E}[\xi_k]=p,

the Weak Law of Large Numbers gives

Snn→Pp.\frac{S_n}{n} \xrightarrow{\mathbb{P}} p.

Thus, for a large number of independent trials, the observed relative frequency of success is close to the probability of success.

This provides a rigorous connection between probability and relative frequency.

Strong Law of Large Numbers

There is a stronger form of the Law of Large Numbers.

The distinction concerns the mode of convergence.

The Weak Law gives convergence in probability:

ξ‾n→Pμ.\overline{\xi}_n\xrightarrow{\mathbb{P}}\mu.

The Strong Law gives almost sure convergence:

P({lim⁡n→∞ξ‾n=μ})=1.\mathbb{P} \left( \left\{ \lim_{n\to\infty}\overline{\xi}_n=\mu \right\} \right) = 1.

We write

ξ‾n→a.s.μ.\overline{\xi}_n\xrightarrow{\text{a.s.}}\mu.

The Strong Law is a deeper result than the Weak Law. Its proof requires tools beyond those developed in this course, so we will use the result without proving it.

For practical applications, both versions express the same fundamental principle: sufficiently many independent observations reveal the underlying mean.

Why the Law of Large Numbers Is Not the Whole Story

The Law of Large Numbers tells us that

ξ‾n≈μ\overline{\xi}_n\approx\mu

for large nn.

But it does not tell us how the random error

ξ‾n−μ\overline{\xi}_n-\mu

behaves.

From Lecture 6,

SD⁡(ξ‾n)=σn.\operatorname{SD}(\overline{\xi}_n) = \frac{\sigma}{\sqrt{n}}.

This suggests that the natural scale of the fluctuations is

1n.\frac{1}{\sqrt{n}}.

Therefore, instead of studying ξ‾n−μ\overline{\xi}_n-\mu directly, we consider the rescaled quantity

Zn=ξ‾n−μσ/n=n(ξ‾n−μ)σ.Z_n = \frac{\overline{\xi}_n-\mu}{\sigma/\sqrt{n}} = \frac{\sqrt{n}(\overline{\xi}_n-\mu)}{\sigma}.

Using Sn=nξ‾nS_n=n\overline{\xi}_n, this can also be written as

Zn=Sn−nμσn.Z_n = \frac{S_n-n\mu}{\sigma\sqrt{n}}.

The Central Limit Theorem describes the limiting distribution of ZnZ_n.

Convergence in Distribution

Before stating the theorem, we introduce another form of convergence.

Let FnF_n be the distribution function of ξn\xi_n, and let FF be the distribution function of ξ\xi.

We say that ξn\xi_n converges in distribution to ξ\xi if

lim⁡n→∞Fn(x)=F(x)\lim_{n\to\infty}F_n(x)=F(x)

at every point xx where FF is continuous.

We write

ξn→dξ.\xi_n\xrightarrow{d}\xi.

Convergence in distribution concerns the limiting shape of the probability distributions. It does not necessarily mean that the random variables themselves become close for individual outcomes.

The Central Limit Theorem

We can now state one of the central results of probability theory.

The theorem says that, after centering by the mean and scaling by the standard deviation, the distribution of the sum approaches a standard normal distribution.

Equivalently, for large nn,

ξ‾n≈N(μ,σ2n).\overline{\xi}_n \approx N\left( \mu,\frac{\sigma^2}{n} \right).

This approximation is one of the main reasons why the normal distribution appears throughout statistics, experimental science, engineering, and numerical analysis.

Interpreting the Central Limit Theorem

The Law of Large Numbers and the Central Limit Theorem answer different questions.

The Law of Large Numbers says

ξ‾n⟶μ,\overline{\xi}_n\longrightarrow\mu,

in an appropriate sense.

The Central Limit Theorem says that the fluctuations around μ\mu occur on the scale 1/n1/\sqrt{n}:

ξ‾n−μ≈σnZ,Z∼N(0,1).\overline{\xi}_n-\mu \approx \frac{\sigma}{\sqrt{n}}Z, \qquad Z\sim N(0,1).

Thus, for large nn,

ξ‾n≈μ+σnZ.\overline{\xi}_n \approx \mu+\frac{\sigma}{\sqrt{n}}Z.

This formula contains both the Law of Large Numbers and the Central Limit Theorem:

The two theorems therefore describe different aspects of the same asymptotic phenomenon.

Approximate Probabilities for Sample Averages

The Central Limit Theorem can be used to approximate probabilities involving ξ‾n\overline{\xi}_n.

Suppose that a<ba<b. Then, for large nn,

P(a≤ξ‾n≤b)≈P(a−μσ/n≤Z≤b−μσ/n),\mathbb{P}(a\leq\overline{\xi}_n\leq b) \approx \mathbb{P} \left( \frac{a-\mu}{\sigma/\sqrt{n}} \leq Z \leq \frac{b-\mu}{\sigma/\sqrt{n}} \right),

where Z∼N(0,1)Z\sim N(0,1).

If Φ\Phi denotes the standard normal distribution function, then

P(a≤ξ‾n≤b)≈Φ((b−μ)nσ)−Φ((a−μ)nσ).\mathbb{P}(a\leq\overline{\xi}_n\leq b) \approx \Phi \left( \frac{(b-\mu)\sqrt{n}}{\sigma} \right) - \Phi \left( \frac{(a-\mu)\sqrt{n}}{\sigma} \right).

This gives a practical way to compute probabilities involving averages even when the original distribution of ξk\xi_k is not normal.

The De Moivre--Laplace Theorem

The normal approximation to the binomial distribution introduced in Lecture 5 is a special case of the Central Limit Theorem.

Let

ξ1,…,ξn\xi_1,\ldots,\xi_n

be independent Bernoulli random variables with parameter pp.

Then

E[ξk]=p,Var⁡(ξk)=p(1−p).\mathbb{E}[\xi_k]=p, \qquad \operatorname{Var}(\xi_k)=p(1-p).

The sum

Sn=∑k=1nξkS_n=\sum_{k=1}^n\xi_k

has a binomial distribution:

Sn∼Bin⁡(n,p).S_n\sim\operatorname{Bin}(n,p).

Applying the Central Limit Theorem gives

Sn−npnp(1−p)→dN(0,1).\frac{S_n-np} {\sqrt{np(1-p)}} \xrightarrow{d} N(0,1).

This is the De Moivre--Laplace theorem.

For large nn, therefore,

Sn≈N(np,np(1−p)).S_n \approx N\left(np,np(1-p)\right).

This explains the normal approximation to the binomial distribution discussed previously.

Continuity correction

The binomial random variable is discrete, whereas the normal distribution is continuous.

When approximating a binomial probability using a normal distribution, a continuity correction can improve the approximation.

For example,

P(Sn≤k)≈Φ(k+12−npnp(1−p)).\mathbb{P}(S_n\leq k) \approx \Phi \left( \frac{k+\frac12-np} {\sqrt{np(1-p)}} \right).

Similarly,

P(Sn=k)≈Φ(k+12−npnp(1−p))−Φ(k−12−npnp(1−p)).\mathbb{P}(S_n=k) \approx \Phi \left( \frac{k+\frac12-np} {\sqrt{np(1-p)}} \right) - \Phi \left( \frac{k-\frac12-np} {\sqrt{np(1-p)}} \right).

The correction accounts for the fact that the integer value kk represents the interval from k−12k-\frac12 to k+12k+\frac12 on the continuous scale.

Numerical Illustration of the Central Limit Theorem

The Central Limit Theorem is remarkable because the original distribution can be very different from the normal distribution.

For example, consider a Bernoulli random variable with

p=0.3.p=0.3.

For each nn, generate independent random variables

ξ1,…,ξn\xi_1,\ldots,\xi_n

and compute

Zn=n(ξ‾n−p)p(1−p).Z_n = \frac{ \sqrt{n}(\overline{\xi}_n-p) }{ \sqrt{p(1-p)} }.

The Central Limit Theorem predicts that the distribution of ZnZ_n approaches the standard normal distribution as nn increases.

import numpy as np
import matplotlib.pyplot as plt

rng = np.random.default_rng(12345)

p = 0.3
n_values = [1, 5, 20, 100]
N = 100000

fig, axes = plt.subplots(2, 2, figsize=(10, 7))

for ax, n in zip(axes.flat, n_values):
    samples = rng.binomial(n, p, size=N)
    averages = samples / n

    z = np.sqrt(n) * (averages - p) / np.sqrt(p * (1 - p))

    ax.hist(z, bins=50, density=True)
    ax.set_title(f"$n={n}$")
    ax.set_xlabel("$Z_n$")
    ax.set_ylabel("Density")

plt.tight_layout()
plt.show()
<Figure size 1000x700 with 4 Axes>

For small nn, the distribution is visibly discrete and far from Gaussian. As nn increases, the standardized distribution becomes increasingly close to the standard normal distribution.

Why Does the Normal Distribution Appear?

The Central Limit Theorem is remarkable because the limiting normal distribution does not depend on the detailed shape of the original distribution.

Suppose that

ξ1,ξ2,…\xi_1,\xi_2,\ldots

are independent and identically distributed random variables with finite mean μ\mu and finite, positive variance σ2\sigma^2.

The individual random variables may have very different distributions. They may be:

Nevertheless, after summation, centering, and normalization, the same limiting distribution appears:

Sn−nμσn→dN(0,1).\frac{S_n-n\mu}{\sigma\sqrt{n}} \xrightarrow{d} N(0,1).

Thus, the detailed distribution of the individual observations becomes less important when many independent contributions are combined.

Schematically,

many independent contributions⟹approximately Gaussian fluctuations.\text{many independent contributions} \quad\Longrightarrow\quad \text{approximately Gaussian fluctuations}.

This is one of the most important universality principles in probability.

The assumptions of the theorem are nevertheless essential. If the random variables are strongly dependent, or if their distributions have sufficiently heavy tails so that the variance is not finite, a different limiting behavior may occur.

An Important Generalization

The classical Central Limit Theorem assumes that the random variables are identically distributed. This assumption can be relaxed.

Suppose that ξ1,ξ2,…\xi_1,\xi_2,\ldots are independent, but not necessarily identically distributed. Define

ak=E[ξk],σk2=Var⁡(ξk),a_k=\mathbb{E}[\xi_k], \qquad \sigma_k^2=\operatorname{Var}(\xi_k),

and

Sn=∑k=1nξk,An=∑k=1nak,Bn2=∑k=1nσk2.S_n=\sum_{k=1}^n\xi_k, \qquad A_n=\sum_{k=1}^n a_k, \qquad B_n^2=\sum_{k=1}^n\sigma_k^2.

Under suitable additional assumptions controlling the contribution of individual random variables, one obtains a result of the form

Sn−AnBn→dN(0,1).\frac{S_n-A_n}{B_n} \xrightarrow{d} N(0,1).

One classical sufficient condition is the Lyapunov condition. If, for some δ>0\delta>0,

∑k=1nE[∣ξk−ak∣2+δ]Bn2+δ⟶0,\frac{ \displaystyle \sum_{k=1}^n \mathbb{E} \left[ |\xi_k-a_k|^{2+\delta} \right] }{ B_n^{2+\delta} } \longrightarrow 0,

then the convergence in (35) holds.

The main idea is that no single random variable should dominate the total fluctuation.

We will not prove this more general version here. It is useful to know that the classical Central Limit Theorem is part of a broader family of limit theorems for sums of independent random variables.

Characteristic Functions

There is another useful way to describe a probability distribution.

For a random variable ξ\xi, its characteristic function is defined by

φξ(t)=E[eitξ],t∈R.\varphi_\xi(t) = \mathbb{E} \left[ e^{it\xi} \right], \qquad t\in\mathbb{R}.

Since

∣eitξ∣=1,|e^{it\xi}|=1,

the characteristic function always exists, even when some moments of ξ\xi do not exist.

Characteristic functions are particularly useful for sums of independent random variables.

If ξ\xi and η\eta are independent, then

φξ+η(t)=φξ(t)φη(t).\varphi_{\xi+\eta}(t) = \varphi_\xi(t)\varphi_\eta(t).

Indeed,

φξ+η(t)=E[eit(ξ+η)]=E[eitξeitη]=E[eitξ]E[eitη]=φξ(t)φη(t).\begin{aligned} \varphi_{\xi+\eta}(t) &= \mathbb{E} \left[ e^{it(\xi+\eta)} \right]\\ &= \mathbb{E} \left[ e^{it\xi}e^{it\eta} \right]\\ &= \mathbb{E}[e^{it\xi}] \mathbb{E}[e^{it\eta}]\\ &= \varphi_\xi(t)\varphi_\eta(t). \end{aligned}

Therefore, if ξ1,…,ξn\xi_1,\ldots,\xi_n are independent,

φSn(t)=∏k=1nφξk(t).\varphi_{S_n}(t) = \prod_{k=1}^n\varphi_{\xi_k}(t).

In the identically distributed case,

φSn(t)=[φξ(t)]n.\varphi_{S_n}(t) = \left[\varphi_\xi(t)\right]^n.

This multiplicative property makes characteristic functions particularly well suited to the study of sums and limit theorems.

For the purposes of this introductory course, it is enough to know the definition and the product property (38). A systematic study of characteristic functions and their role in proving limit theorems belongs to a more advanced treatment of probability theory.

Law of Large Numbers and Central Limit Theorem

The two main results of this lecture describe different aspects of the behavior of averages.

The Law of Large Numbers states that

ξ‾n→Pμ.\overline{\xi}_n \xrightarrow{\mathbb{P}} \mu.

Thus, as the number of observations increases, the average approaches the expected value.

The Central Limit Theorem describes the fluctuations around this limiting value:

n(ξ‾n−μ)σ→dN(0,1).\frac{ \sqrt{n}(\overline{\xi}_n-\mu) }{ \sigma } \xrightarrow{d} N(0,1).

Thus, the typical size of the error

ξ‾n−μ\overline{\xi}_n-\mu

is of order 1/n1/\sqrt{n}, and the appropriately rescaled error becomes approximately Gaussian.

The two results can therefore be summarized schematically as

ξ‾n≈μ+σnZ,Z∼N(0,1)\boxed{ \overline{\xi}_n \approx \mu+\frac{\sigma}{\sqrt{n}}Z, \qquad Z\sim N(0,1) }

for large nn.

The Law of Large Numbers concerns the location of the average, while the Central Limit Theorem concerns its fluctuations.

From Probability to Statistics

These results provide the probabilistic foundation for the estimation of unknown quantities.

Suppose that

ξ1,…,ξn\xi_1,\ldots,\xi_n

are independent observations from a distribution with unknown mean μ\mu.

A natural estimator of μ\mu is the sample mean

μ^n=ξ‾n=1n∑k=1nξk.\widehat{\mu}_n = \overline{\xi}_n = \frac{1}{n}\sum_{k=1}^n\xi_k.

The Law of Large Numbers gives

μ^n→Pμ.\widehat{\mu}_n \xrightarrow{\mathbb{P}} \mu.

In other words, the sample mean is a consistent estimator of the population mean.

The Central Limit Theorem gives the asymptotic distribution of the estimation error:

n(μ^n−μ)σ→dN(0,1).\frac{ \sqrt{n}(\widehat{\mu}_n-\mu) }{ \sigma } \xrightarrow{d} N(0,1).

Equivalently, for large nn,

μ^n≈N(μ,σ2n).\widehat{\mu}_n \approx N\left( \mu,\frac{\sigma^2}{n} \right).

This approximation is the starting point for many statistical procedures, including confidence intervals and hypothesis tests.

Final Perspective

The previous lectures have developed the basic language and tools of probability.

We began with random experiments, events, and the axiomatic definition of probability. We then introduced conditional probability and independence, which allow us to describe how information about one event or random quantity affects another.

For a single random variable, we studied:

We then extended the theory to pairs of random variables and introduced:

Finally, by considering sequences of random variables, we studied the asymptotic behavior of averages.

The Law of Large Numbers explains why averages stabilize:

ξ‾n→Pμ.\overline{\xi}_n \xrightarrow{\mathbb{P}} \mu.

The Central Limit Theorem explains the fluctuations around the limiting value:

n(ξ‾n−μ)σ→dN(0,1).\frac{ \sqrt{n}(\overline{\xi}_n-\mu) }{ \sigma } \xrightarrow{d} N(0,1).

These two results provide the fundamental bridge from probability to statistics.

They explain why repeated observations can be used to estimate unknown quantities, why averaging reduces random fluctuations, and why the normal distribution appears so frequently when many independent random effects are combined.

At the same time, the theory developed in these lectures provides the foundation for many applications:

The central message can be summarized as follows:

individual random observations→averagingstable macroscopic behavior\boxed{ \text{individual random observations} \quad \xrightarrow{\text{averaging}} \quad \text{stable macroscopic behavior} }

with the Law of Large Numbers describing the limiting value and the Central Limit Theorem describing the fluctuations around it.