Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Lecture 7 - Joint Distributions and Joint Moments

Joint Distributions and Joint Moments

So far, we have studied random variables one at a time. We have introduced their probability distributions, cumulative distribution functions, densities, expectations, moments, and variances.

In many applications, however, several random quantities are observed simultaneously.

Examples include:

In such situations, knowing the distribution of each random variable separately is generally not enough. We also need to describe their joint behavior.

This leads to the concepts of:


The case of two random variables

Let

ξ,η:Ω→R\xi,\eta:\Omega\to\mathbb{R}

be two random variables defined on the same probability space.

The pair

(ξ,η)(\xi,\eta)

can itself be regarded as a random variable taking values in R2\mathbb{R}^2.

For example, if ξ\xi represents the temperature of a system and η\eta represents its pressure, then one observation of the experiment produces a point

(ξ(ω),η(ω))∈R2.(\xi(\omega),\eta(\omega)) \in\mathbb{R}^2.

The distribution of this random point describes how the two quantities behave together.


Joint distribution function

The joint cumulative distribution function of ξ\xi and η\eta is defined by

Fξ,η(x,y)=P({ξ≤x,η≤y}).F_{\xi,\eta}(x,y) = \mathbb{P} \left( \{\xi\leq x,\eta\leq y\} \right).

This function gives the probability that the random point (ξ,η)(\xi,\eta) lies in the lower-left rectangle

(−∞,x]×(−∞,y].(-\infty,x]\times(-\infty,y].

The joint CDF contains the complete information about the joint distribution of the pair (ξ,η)(\xi,\eta).

In particular, it can be used to compute the probability of a rectangle:

P({a<ξ≤b, c<η≤d})\mathbb{P} \left( \{a<\xi\leq b,\ c<\eta\leq d\} \right)

by combining four values of the joint CDF:

P({a<ξ≤b, c<η≤d})=Fξ,η(b,d)−Fξ,η(a,d)−Fξ,η(b,c)+Fξ,η(a,c).\begin{aligned} &\mathbb{P} \left( \{a<\xi\leq b,\ c<\eta\leq d\} \right)\\ &\qquad= F_{\xi,\eta}(b,d) -F_{\xi,\eta}(a,d) -F_{\xi,\eta}(b,c) +F_{\xi,\eta}(a,c). \end{aligned}

This is the two-dimensional analogue of the relation between a one-dimensional CDF and interval probabilities.


Joint distribution of discrete random variables

Suppose that ξ\xi and η\eta are discrete random variables.

Their joint probability mass function is

Pξ,η(x,y)=P({ξ=x,η=y}).P_{\xi,\eta}(x,y) = \mathbb{P} \left( \{\xi=x,\eta=y\} \right).

The probabilities must satisfy

Pξ,η(x,y)≥0P_{\xi,\eta}(x,y)\geq0

and

∑x∑yPξ,η(x,y)=1.\sum_x\sum_y P_{\xi,\eta}(x,y) = 1.

The joint mass function gives the probability of every possible pair of values.

For example, suppose that two components are tested and

ξ={1,first component is defective,0,otherwise,\xi= \begin{cases} 1,&\text{first component is defective},\\ 0,&\text{otherwise}, \end{cases}

and similarly for η\eta.

The joint distribution distinguishes between the probabilities

P({ξ=1,η=1}),P({ξ=1,η=0}),P({ξ=0,η=1}),P({ξ=0,η=0}).\begin{split} \mathbb{P}(\{\xi=1,\eta=1\}), \qquad \mathbb{P}(\{\xi=1,\eta=0\}), \\ \mathbb{P}(\{\xi=0,\eta=1\}), \qquad \mathbb{P}(\{\xi=0,\eta=0\}). \end{split}

These four probabilities describe how the two defect indicators behave together.


Marginal distributions

Suppose that ξ\xi and η\eta are discrete random variables with joint mass function Pξ,ηP_{\xi,\eta}.

The distribution of ξ\xi alone is obtained by summing over all possible values of η\eta:

Pξ(x)=∑yPξ,η(x,y).P_\xi(x) = \sum_y P_{\xi,\eta}(x,y).

Similarly,

Pη(y)=∑xPξ,η(x,y).P_\eta(y) = \sum_x P_{\xi,\eta}(x,y).

These are called the marginal distributions.

The terminology reflects the fact that a joint probability table can be displayed with the marginal probabilities in its margins.


Continuous joint distributions

Suppose that (ξ,η)(\xi,\eta) is a pair of continuous random variables.

A joint probability density function is a nonnegative function

pξ,η(x,y)p_{\xi,\eta}(x,y)

satisfying

∫−∞+∞∫−∞+∞pξ,η(x,y) dx dy=1.\int_{-\infty}^{+\infty} \int_{-\infty}^{+\infty} p_{\xi,\eta}(x,y)\,dx\,dy = 1.

The probability that (ξ,η)(\xi,\eta) belongs to a region D⊆R2D\subseteq\mathbb{R}^2 is

P({(ξ,η)∈D})=∬Dpξ,η(x,y) dx dy.\mathbb{P} \left( \{(\xi,\eta)\in D\} \right) = \iint_D p_{\xi,\eta}(x,y)\,dx\,dy.

For a rectangle,

P({a≤ξ≤b, c≤η≤d})=∫ab∫cdpξ,η(x,y) dy dx.\begin{aligned} &\mathbb{P} \left( \{a\leq\xi\leq b,\ c\leq\eta\leq d\} \right)\\ &\qquad= \int_a^b\int_c^d p_{\xi,\eta}(x,y)\,dy\,dx. \end{aligned}

Thus, just as in one dimension, probabilities correspond to areas under a density. In two dimensions, the density is a surface and probabilities correspond to volumes under that surface.


Marginal densities

The marginal density of ξ\xi is obtained by integrating out η\eta:

pξ(x)=∫−∞+∞pξ,η(x,y) dy.p_\xi(x) = \int_{-\infty}^{+\infty} p_{\xi,\eta}(x,y)\,dy.

Similarly,

pη(y)=∫−∞+∞pξ,η(x,y) dx.p_\eta(y) = \int_{-\infty}^{+\infty} p_{\xi,\eta}(x,y)\,dx.

This is the continuous analogue of summing the rows and columns of a discrete joint probability table.

The terminology “marginal” is therefore used in both cases:

Pξ(x)=∑yPξ,η(x,y),pξ(x)=∫pξ,η(x,y) dy,Pη(y)=∑xPξ,η(x,y),pη(y)=∫pξ,η(x,y) dx.\boxed{ \begin{aligned} P_\xi(x) &= \sum_y P_{\xi,\eta}(x,y), & p_\xi(x) &= \int p_{\xi,\eta}(x,y)\,dy, \\[6pt] P_\eta(y) &= \sum_x P_{\xi,\eta}(x,y), & p_\eta(y) &= \int p_{\xi,\eta}(x,y)\,dx. \end{aligned}}

Conditional Distributions

The joint distribution describes how two random variables behave together, while the marginal distributions describe each variable separately. A third useful description is obtained when the value of one random variable is known.

Suppose that ξ\xi and η\eta are random variables. If we know that η=y\eta=y, we may ask how the distribution of ξ\xi changes. This leads to the notion of a conditional distribution.

Conditional distribution in the discrete case

Suppose that ξ\xi and η\eta are discrete random variables. For a value yy such that

Pη(y)>0,P_\eta(y)>0,

the conditional probability that ξ=x\xi=x given η=y\eta=y is

Pξ∣η(x∣y)=P({ξ=x}∣{η=y})=Pξ,η(x,y)Pη(y).P_{\xi\mid\eta}(x\mid y) = \mathbb{P}(\{\xi=x\}\mid\{\eta=y\}) = \frac{P_{\xi,\eta}(x,y)}{P_\eta(y)}.

For every fixed yy with Pη(y)>0P_\eta(y)>0, this is a probability distribution in xx. In particular,

Pξ∣η(x∣y)≥0,P_{\xi\mid\eta}(x\mid y)\geq 0,

and

∑xPξ∣η(x∣y)=1.\sum_x P_{\xi\mid\eta}(x\mid y)=1.

The joint distribution can therefore be reconstructed from the marginal and conditional distributions:

Pξ,η(x,y)=Pξ∣η(x∣y)Pη(y).P_{\xi,\eta}(x,y) = P_{\xi\mid\eta}(x\mid y)P_\eta(y).

This is the random-variable version of the multiplication rule introduced for events in Lecture 2.

Conditional density in the continuous case

Suppose now that the pair (ξ,η)(\xi,\eta) admits a joint density pξ,ηp_{\xi,\eta}, and let pηp_\eta be the marginal density of η\eta.

For values yy such that pη(y)>0p_\eta(y)>0, the conditional density of ξ\xi given η=y\eta=y is defined by

pξ∣η(x∣y)=pξ,η(x,y)pη(y).p_{\xi\mid\eta}(x\mid y) = \frac{p_{\xi,\eta}(x,y)}{p_\eta(y)}.

For each fixed yy for which the conditional density is defined,

∫−∞+∞pξ∣η(x∣y) dx=1.\int_{-\infty}^{+\infty} p_{\xi\mid\eta}(x\mid y)\,dx = 1.

The corresponding conditional distribution function is

Fξ∣η(x∣y)=∫−∞xpξ∣η(t∣y) dt.F_{\xi\mid\eta}(x\mid y) = \int_{-\infty}^{x} p_{\xi\mid\eta}(t\mid y)\,dt.

As in the discrete case, the joint density can be written as

pξ,η(x,y)=pξ∣η(x∣y)pη(y).p_{\xi,\eta}(x,y) = p_{\xi\mid\eta}(x\mid y)p_\eta(y).

Independence of Random Variables

For events, independence means that knowing whether one event occurred does not change the probability of another event. The same idea extends naturally to random variables.

Two random variables ξ\xi and η\eta are said to be independent if, for every pair of Borel sets A,B⊆RA,B\subseteq\mathbb{R},

P({ξ∈A}∩{η∈B})=P({ξ∈A})P({η∈B}).\mathbb{P}(\{\xi\in A\}\cap\{\eta\in B\}) = \mathbb{P}(\{\xi\in A\}) \mathbb{P}(\{\eta\in B\}).

In words, observing the value of ξ\xi provides no probabilistic information about the value of η\eta, and vice versa.

Independence in the discrete case

For discrete random variables, independence is equivalent to the factorization

Pξ,η(x,y)=Pξ(x)Pη(y)P_{\xi,\eta}(x,y) = P_\xi(x)P_\eta(y)

for every pair (x,y)(x,y).

For example, if ξ\xi and η\eta are independent Bernoulli random variables with parameters pp and qq, respectively, then

P({ξ=1,η=1})=pq,P({ξ=1,η=0})=p(1−q),P({ξ=0,η=1})=(1−p)q,P({ξ=0,η=0})=(1−p)(1−q).\begin{aligned} \mathbb{P}(\{\xi=1,\eta=1\}) &= pq,\\ \mathbb{P}(\{\xi=1,\eta=0\}) &= p(1-q),\\ \mathbb{P}(\{\xi=0,\eta=1\}) &= (1-p)q,\\ \mathbb{P}(\{\xi=0,\eta=0\}) &= (1-p)(1-q). \end{aligned}

The joint distribution is completely determined by the two marginal distributions.

Independence in the continuous case

If (ξ,η)(\xi,\eta) admits a joint density, independence is equivalent to

pξ,η(x,y)=pξ(x)pη(y)p_{\xi,\eta}(x,y) = p_\xi(x)p_\eta(y)

for almost every (x,y)∈R2(x,y)\in\mathbb{R}^2.

Equivalently, the joint distribution function factorizes as

Fξ,η(x,y)=Fξ(x)Fη(y).F_{\xi,\eta}(x,y) = F_\xi(x)F_\eta(y).

Thus, independence is a special situation in which the joint distribution contains no additional information beyond the two marginal distributions.

Independence and conditional distributions

The connection with conditional distributions is particularly useful.

If ξ\xi and η\eta are independent, then, whenever the conditional distribution is defined,

Pξ∣η(x∣y)=Pξ(x)P_{\xi\mid\eta}(x\mid y) = P_\xi(x)

in the discrete case, and

pξ∣η(x∣y)=pξ(x)p_{\xi\mid\eta}(x\mid y) = p_\xi(x)

in the continuous case.

Thus, conditioning on η\eta does not change the distribution of ξ\xi.

The converse also holds under the usual conditions: if the conditional distribution of ξ\xi does not depend on yy, then ξ\xi and η\eta are independent.

Functions of Two Random Variables

Once a joint distribution is available, we can compute expectations not only of ξ\xi and η\eta separately, but also of functions involving both variables.

Let

g:R2→Rg:\mathbb{R}^2\to\mathbb{R}

be a suitable function.

If ξ\xi and η\eta are discrete, then

E[g(ξ,η)]=∑x∑yg(x,y)Pξ,η(x,y).\mathbb{E}[g(\xi,\eta)] = \sum_x\sum_y g(x,y)P_{\xi,\eta}(x,y).

If (ξ,η)(\xi,\eta) admits a joint density, then

E[g(ξ,η)]=∫−∞+∞∫−∞+∞g(x,y)pξ,η(x,y) dx dy.\mathbb{E}[g(\xi,\eta)] = \int_{-\infty}^{+\infty} \int_{-\infty}^{+\infty} g(x,y)p_{\xi,\eta}(x,y)\,dx\,dy.

These formulas are direct extensions of the one-dimensional formulas introduced in Lecture 6.

For example, choosing

g(x,y)=xyg(x,y)=xy

gives the mixed first moment

E[ξη].\mathbb{E}[\xi\eta].

More generally, the quantities

E[ξrηs],r,s∈N,\mathbb{E}[\xi^r\eta^s], \qquad r,s\in\mathbb{N},

are called joint moments of ξ\xi and η\eta.

The exponents determine the order of the moment. For example,

E[ξη]\mathbb{E}[\xi\eta]

is a mixed moment of total order two, while

E[ξ2η]\mathbb{E}[\xi^2\eta]

has total order three.

Products of independent random variables

Independence leads to an important simplification.

If ξ\xi and η\eta are independent and the expectations exist, then

E[ξη]=E[ξ]E[η].\mathbb{E}[\xi\eta] = \mathbb{E}[\xi]\mathbb{E}[\eta].

More generally, if the relevant expectations exist,

E[g(ξ)h(η)]=E[g(ξ)]E[h(η)].\mathbb{E}[g(\xi)h(\eta)] = \mathbb{E}[g(\xi)]\mathbb{E}[h(\eta)].

For the discrete case, this follows immediately from the factorization of the joint pmf:

E[g(ξ)h(η)]=∑x∑yg(x)h(y)Pξ,η(x,y)=∑x∑yg(x)h(y)Pξ(x)Pη(y)=(∑xg(x)Pξ(x))(∑yh(y)Pη(y)).\begin{aligned} \mathbb{E}[g(\xi)h(\eta)] &= \sum_x\sum_y g(x)h(y)P_{\xi,\eta}(x,y)\\ &= \sum_x\sum_y g(x)h(y)P_\xi(x)P_\eta(y)\\ &= \left(\sum_x g(x)P_\xi(x)\right) \left(\sum_y h(y)P_\eta(y)\right). \end{aligned}

The continuous case follows in the same way by using the factorization of the joint density.

Covariance

The product E[ξη]\mathbb{E}[\xi\eta] alone does not tell us whether large values of ξ\xi tend to occur together with large values of η\eta. To measure this type of joint variation, we introduce the covariance.

Assume that ξ\xi and η\eta have finite second moments. Their covariance is defined by

Cov⁡(ξ,η)=E[(ξ−E[ξ])(η−E[η])].\operatorname{Cov}(\xi,\eta) = \mathbb{E} \left[ (\xi-\mathbb{E}[\xi]) (\eta-\mathbb{E}[\eta]) \right].

The covariance measures whether the two variables tend to deviate from their means in the same direction.

Expanding the product gives

Cov⁡(ξ,η)=E[ξη]−E[ξ]E[η].\operatorname{Cov}(\xi,\eta) = \mathbb{E}[\xi\eta] - \mathbb{E}[\xi]\mathbb{E}[\eta].

Thus, covariance compares the mixed moment E[ξη]\mathbb{E}[\xi\eta] with the product of the means.

If large values of ξ\xi tend to be associated with large values of η\eta, the covariance is typically positive. If large values of one variable tend to be associated with small values of the other, it is typically negative.

The covariance is symmetric:

Cov⁡(ξ,η)=Cov⁡(η,ξ).\operatorname{Cov}(\xi,\eta) = \operatorname{Cov}(\eta,\xi).

It is also linear in each argument. For constants a,b,c,da,b,c,d,

Cov⁡(aξ+b,cη+d)=ac Cov⁡(ξ,η).\operatorname{Cov}(a\xi+b,c\eta+d) = ac\,\operatorname{Cov}(\xi,\eta).

In particular,

Cov⁡(ξ,c)=0\operatorname{Cov}(\xi,c)=0

for every constant cc.

Most importantly, variance is a special case of covariance:

Cov⁡(ξ,ξ)=Var⁡(ξ).\operatorname{Cov}(\xi,\xi) = \operatorname{Var}(\xi).

Variance of a sum

Using covariance, the variance of a sum can be written as

Var⁡(ξ+η)=Var⁡(ξ)+Var⁡(η)+2Cov⁡(ξ,η).\operatorname{Var}(\xi+\eta) = \operatorname{Var}(\xi) + \operatorname{Var}(\eta) + 2\operatorname{Cov}(\xi,\eta).

More generally,

Var⁡(∑k=1nξk)=∑k=1nVar⁡(ξk)+2∑1≤i<j≤nCov⁡(ξi,ξj).\operatorname{Var} \left( \sum_{k=1}^n \xi_k \right) = \sum_{k=1}^n\operatorname{Var}(\xi_k) + 2\sum_{1\leq i<j\leq n} \operatorname{Cov}(\xi_i,\xi_j).

If the random variables are pairwise uncorrelated, all covariance terms vanish and therefore

Var⁡(∑k=1nξk)=∑k=1nVar⁡(ξk).\operatorname{Var} \left( \sum_{k=1}^n \xi_k \right) = \sum_{k=1}^n\operatorname{Var}(\xi_k).

In particular, independent random variables are uncorrelated whenever their second moments exist. Hence the variance formula for sums of independent random variables introduced in Lecture 6 is a consequence of the more general covariance formula.

Correlation

The numerical value of covariance depends on the units in which the variables are measured. For example, changing a measurement from metres to centimetres multiplies the covariance by a factor of 102.

To obtain a dimensionless measure, we normalize the covariance by the standard deviations.

Suppose that

0<Var⁡(ξ)<∞,0<Var⁡(η)<∞.0<\operatorname{Var}(\xi)<\infty, \qquad 0<\operatorname{Var}(\eta)<\infty.

The correlation coefficient of ξ\xi and η\eta is

ρξ,η=Cov⁡(ξ,η)Var⁡(ξ)Var⁡(η).\rho_{\xi,\eta} = \frac{\operatorname{Cov}(\xi,\eta)} {\sqrt{\operatorname{Var}(\xi)} \sqrt{\operatorname{Var}(\eta)}}.

Using the standard deviations σξ\sigma_\xi and ση\sigma_\eta, this can be written as

ρξ,η=Cov⁡(ξ,η)σξση.\rho_{\xi,\eta} = \frac{\operatorname{Cov}(\xi,\eta)} {\sigma_\xi\sigma_\eta}.

The Cauchy--Schwarz inequality implies that

−1≤ρξ,η≤1.-1\leq \rho_{\xi,\eta}\leq 1.

The value of ρξ,η\rho_{\xi,\eta} describes the strength and direction of a linear association between the two variables:

It is important to distinguish correlation from independence.

ξ and η independent⟹ρξ,η=0,\xi\text{ and }\eta\text{ independent} \quad\Longrightarrow\quad \rho_{\xi,\eta}=0,

provided the variances are finite and positive.

The converse is generally false.

Zero correlation does not imply independence

Consider a random variable

ξ∼U(−1,1)\xi\sim U(-1,1)

and define

η=ξ2.\eta=\xi^2.

Clearly, η\eta is completely determined by ξ\xi, so the two variables cannot be independent.

However, by symmetry,

E[ξ]=0\mathbb{E}[\xi]=0

and

E[ξ3]=0.\mathbb{E}[\xi^3]=0.

Consequently,

Cov⁡(ξ,η)=Cov⁡(ξ,ξ2)=E[ξ3]−E[ξ]E[ξ2]=0.\operatorname{Cov}(\xi,\eta) = \operatorname{Cov}(\xi,\xi^2) = \mathbb{E}[\xi^3] - \mathbb{E}[\xi]\mathbb{E}[\xi^2] = 0.

Thus ξ\xi and η\eta are uncorrelated but not independent.

An Interpretation Through Linear Prediction

The correlation coefficient also has a useful interpretation in terms of linear prediction.

Suppose we want to approximate ξ\xi by an affine function of η\eta,

ξ^=c1+c2η.\widehat{\xi}=c_1+c_2\eta.

We choose c1c_1 and c2c_2 to minimize the mean squared error

E[(ξ−c1−c2η)2].\mathbb{E} \left[ (\xi-c_1-c_2\eta)^2 \right].

Assume that ξ\xi and η\eta have finite, positive variances. The optimal coefficients are

c2=Cov⁡(ξ,η)Var⁡(η)c_2 = \frac{\operatorname{Cov}(\xi,\eta)} {\operatorname{Var}(\eta)}

and

c1=E[ξ]−c2E[η].c_1 = \mathbb{E}[\xi] - c_2\mathbb{E}[\eta].

The corresponding minimum mean squared error is

min⁡c1,c2E[(ξ−c1−c2η)2]=Var⁡(ξ)(1−ρξ,η2).\min_{c_1,c_2} \mathbb{E} \left[ (\xi-c_1-c_2\eta)^2 \right] = \operatorname{Var}(\xi)(1-\rho_{\xi,\eta}^2).

Thus ρξ,η2\rho_{\xi,\eta}^2 measures the fraction of the variance of ξ\xi that can be explained by the best affine prediction based on η\eta.

This interpretation also clarifies why zero correlation does not necessarily mean independence. If ρξ,η=0\rho_{\xi,\eta}=0, the best affine predictor of ξ\xi based on η\eta is simply its mean. Nevertheless, nonlinear relationships between the variables may still exist, as in the example η=ξ2\eta=\xi^2.

Summary

The joint distribution provides a complete probabilistic description of two random variables and allows us to study their dependence.

For discrete random variables, the joint distribution is described by

Pξ,η(x,y)=P({ξ=x,η=y}),P_{\xi,\eta}(x,y) = \mathbb{P}(\{\xi=x,\eta=y\}),

while for a pair admitting a joint density it is described by pξ,η(x,y)p_{\xi,\eta}(x,y).

From the joint distribution we can obtain:

ObjectDiscrete caseContinuous case
Joint distributionPξ,η(x,y)P_{\xi,\eta}(x,y)pξ,η(x,y)p_{\xi,\eta}(x,y)
Marginal of ξ\xiPξ(x)=∑yPξ,η(x,y)\displaystyle P_\xi(x)=\sum_yP_{\xi,\eta}(x,y)pξ(x)=∫pξ,η(x,y) dy\displaystyle p_\xi(x)=\int p_{\xi,\eta}(x,y)\,dy
Conditional distributionPξ∣η(x∣y)=Pξ,η(x,y)Pη(y)\displaystyle P_{\xi\mid\eta}(x\mid y)=\frac{P_{\xi,\eta}(x,y)}{P_\eta(y)}pξ∣η(x∣y)=pξ,η(x,y)pη(y)\displaystyle p_{\xi\mid\eta}(x\mid y)=\frac{p_{\xi,\eta}(x,y)}{p_\eta(y)}
IndependencePξ,η=PξPηP_{\xi,\eta}=P_\xi P_\etapξ,η=pξpηp_{\xi,\eta}=p_\xi p_\eta
Joint moment∑x∑yxrysPξ,η(x,y)\displaystyle \sum_x\sum_y x^ry^sP_{\xi,\eta}(x,y)∬xryspξ,η(x,y) dx dy\displaystyle \iint x^ry^s p_{\xi,\eta}(x,y)\,dx\,dy
CovarianceCov⁡(ξ,η)=E[ξη]−E[ξ]E[η]\operatorname{Cov}(\xi,\eta)=\mathbb{E}[\xi\eta]-\mathbb{E}[\xi]\mathbb{E}[\eta]same
Correlationρξ,η=Cov⁡(ξ,η)σξση\displaystyle \rho_{\xi,\eta}=\frac{\operatorname{Cov}(\xi,\eta)}{\sigma_\xi\sigma_\eta}same

The main relationships to remember are

joint distribution⟶{marginal distributions,conditional distributions,joint moments.\text{joint distribution} \quad\longrightarrow\quad \begin{cases} \text{marginal distributions},\\ \text{conditional distributions},\\ \text{joint moments}. \end{cases}

Independence is characterized by factorization of the joint distribution:

Pξ,η(x,y)=Pξ(x)Pη(y)P_{\xi,\eta}(x,y)=P_\xi(x)P_\eta(y)

in the discrete case, or

pξ,η(x,y)=pξ(x)pη(y)p_{\xi,\eta}(x,y)=p_\xi(x)p_\eta(y)

in the continuous case.

Finally,

independence⟹Cov⁡(ξ,η)=0,\text{independence} \quad\Longrightarrow\quad \operatorname{Cov}(\xi,\eta)=0,

but zero covariance does not generally imply independence.

From Two Random Variables to Many Observations

The concepts introduced in this lecture provide the language needed to study collections of random variables.

In particular, we can now distinguish between:

These ideas become especially important when considering a sequence of random variables

ξ1,ξ2,…,ξn,…\xi_1,\xi_2,\ldots,\xi_n,\ldots

and quantities such as their sum

Sn=∑k=1nξkS_n=\sum_{k=1}^n\xi_k

or their average

ξ‾n=1n∑k=1nξk.\overline{\xi}_n = \frac{1}{n}\sum_{k=1}^n\xi_k.

The variance calculation from Lecture 6 shows that, for independent identically distributed random variables with variance σ2\sigma^2,

Var⁡(ξ‾n)=σ2n.\operatorname{Var}(\overline{\xi}_n) = \frac{\sigma^2}{n}.

Thus the average becomes increasingly concentrated around its mean as the number of observations grows.

In the next lecture we will study this phenomenon more systematically. The Law of Large Numbers describes the convergence of averages to their expected value, while the Central Limit Theorem describes the fluctuations around that limiting value and explains the fundamental role of the normal distribution in statistics and scientific computation.