Lecture 7 - Joint Distributions and Joint Moments
Joint Distributions and Joint Moments ¶ So far, we have studied random variables one at a time. We have introduced their probability distributions, cumulative distribution functions, densities, expectations, moments, and variances.
In many applications, however, several random quantities are observed simultaneously.
Examples include:
the position and velocity of a particle;
the temperature and pressure of a physical system;
the errors of two measuring instruments;
the dimensions of two components produced by the same manufacturing process;
the number of events observed in two different regions.
In such situations, knowing the distribution of each random variable separately is generally not enough. We also need to describe their joint behavior .
This leads to the concepts of:
joint distributions;
marginal distributions;
conditional distributions;
independence of random variables;
joint moments;
covariance and correlation.
The case of two random variables ¶ Let
ξ , η : Ω → R \xi,\eta:\Omega\to\mathbb{R} ξ , η : Ω → R be two random variables defined on the same probability space.
The pair
can itself be regarded as a random variable taking values in R 2 \mathbb{R}^2 R 2 .
For example, if ξ \xi ξ represents the temperature of a system and η \eta η represents its pressure, then one observation of the experiment produces a point
( ξ ( ω ) , η ( ω ) ) ∈ R 2 . (\xi(\omega),\eta(\omega))
\in\mathbb{R}^2. ( ξ ( ω ) , η ( ω )) ∈ R 2 . The distribution of this random point describes how the two quantities behave together.
Joint distribution function ¶ The joint cumulative distribution function of ξ \xi ξ and η \eta η is defined by
F ξ , η ( x , y ) = P ( { ξ ≤ x , η ≤ y } ) . F_{\xi,\eta}(x,y)
=
\mathbb{P}
\left(
\{\xi\leq x,\eta\leq y\}
\right). F ξ , η ( x , y ) = P ( { ξ ≤ x , η ≤ y } ) . This function gives the probability that the random point ( ξ , η ) (\xi,\eta) ( ξ , η ) lies in the lower-left rectangle
( − ∞ , x ] × ( − ∞ , y ] . (-\infty,x]\times(-\infty,y]. ( − ∞ , x ] × ( − ∞ , y ] . The joint CDF contains the complete information about the joint distribution of the pair ( ξ , η ) (\xi,\eta) ( ξ , η ) .
In particular, it can be used to compute the probability of a rectangle:
P ( { a < ξ ≤ b , c < η ≤ d } ) \mathbb{P}
\left(
\{a<\xi\leq b,\ c<\eta\leq d\}
\right) P ( { a < ξ ≤ b , c < η ≤ d } ) by combining four values of the joint CDF:
P ( { a < ξ ≤ b , c < η ≤ d } ) = F ξ , η ( b , d ) − F ξ , η ( a , d ) − F ξ , η ( b , c ) + F ξ , η ( a , c ) . \begin{aligned}
&\mathbb{P}
\left(
\{a<\xi\leq b,\ c<\eta\leq d\}
\right)\\
&\qquad=
F_{\xi,\eta}(b,d)
-F_{\xi,\eta}(a,d)
-F_{\xi,\eta}(b,c)
+F_{\xi,\eta}(a,c).
\end{aligned} P ( { a < ξ ≤ b , c < η ≤ d } ) = F ξ , η ( b , d ) − F ξ , η ( a , d ) − F ξ , η ( b , c ) + F ξ , η ( a , c ) . This is the two-dimensional analogue of the relation between a one-dimensional CDF and interval probabilities.
Joint distribution of discrete random variables ¶ Suppose that ξ \xi ξ and η \eta η are discrete random variables.
Their joint probability mass function is
P ξ , η ( x , y ) = P ( { ξ = x , η = y } ) . P_{\xi,\eta}(x,y)
=
\mathbb{P}
\left(
\{\xi=x,\eta=y\}
\right). P ξ , η ( x , y ) = P ( { ξ = x , η = y } ) . The probabilities must satisfy
P ξ , η ( x , y ) ≥ 0 P_{\xi,\eta}(x,y)\geq0 P ξ , η ( x , y ) ≥ 0 and
∑ x ∑ y P ξ , η ( x , y ) = 1. \sum_x\sum_y
P_{\xi,\eta}(x,y)
=
1. x ∑ y ∑ P ξ , η ( x , y ) = 1. The joint mass function gives the probability of every possible pair of values.
For example, suppose that two components are tested and
ξ = { 1 , first component is defective , 0 , otherwise , \xi=
\begin{cases}
1,&\text{first component is defective},\\
0,&\text{otherwise},
\end{cases} ξ = { 1 , 0 , first component is defective , otherwise , and similarly for η \eta η .
The joint distribution distinguishes between the probabilities
P ( { ξ = 1 , η = 1 } ) , P ( { ξ = 1 , η = 0 } ) , P ( { ξ = 0 , η = 1 } ) , P ( { ξ = 0 , η = 0 } ) . \begin{split}
\mathbb{P}(\{\xi=1,\eta=1\}),
\qquad
\mathbb{P}(\{\xi=1,\eta=0\}),
\\
\mathbb{P}(\{\xi=0,\eta=1\}),
\qquad
\mathbb{P}(\{\xi=0,\eta=0\}).
\end{split} P ({ ξ = 1 , η = 1 }) , P ({ ξ = 1 , η = 0 }) , P ({ ξ = 0 , η = 1 }) , P ({ ξ = 0 , η = 0 }) . These four probabilities describe how the two defect indicators behave together.
Marginal distributions ¶ Suppose that ξ \xi ξ and η \eta η are discrete random variables with joint mass function P ξ , η P_{\xi,\eta} P ξ , η .
The distribution of ξ \xi ξ alone is obtained by summing over all possible values of η \eta η :
P ξ ( x ) = ∑ y P ξ , η ( x , y ) . P_\xi(x)
=
\sum_y P_{\xi,\eta}(x,y). P ξ ( x ) = y ∑ P ξ , η ( x , y ) . Similarly,
P η ( y ) = ∑ x P ξ , η ( x , y ) . P_\eta(y)
=
\sum_x P_{\xi,\eta}(x,y). P η ( y ) = x ∑ P ξ , η ( x , y ) . These are called the marginal distributions .
The terminology reflects the fact that a joint probability table can be displayed with the marginal probabilities in its margins.
Suppose that the joint distribution of ξ \xi ξ and η \eta η is
η = 0 η = 1 Total ξ = 0 0.50 0.10 0.60 ξ = 1 0.20 0.20 0.40 Total 0.70 0.30 1 \begin{array}{c|cc|c}
& \eta=0 & \eta=1 & \text{Total}\\
\hline
\xi=0 & 0.50 & 0.10 & 0.60\\
\xi=1 & 0.20 & 0.20 & 0.40\\
\hline
\text{Total} & 0.70 & 0.30 & 1
\end{array} ξ = 0 ξ = 1 Total η = 0 0.50 0.20 0.70 η = 1 0.10 0.20 0.30 Total 0.60 0.40 1 Then
P ξ ( 0 ) = 0.60 , P ξ ( 1 ) = 0.40 , P_\xi(0)=0.60,
\qquad
P_\xi(1)=0.40, P ξ ( 0 ) = 0.60 , P ξ ( 1 ) = 0.40 , while
P η ( 0 ) = 0.70 , P η ( 1 ) = 0.30. P_\eta(0)=0.70,
\qquad
P_\eta(1)=0.30. P η ( 0 ) = 0.70 , P η ( 1 ) = 0.30. The marginal distributions describe each variable separately, but they do not contain all the information in the joint distribution.
Continuous joint distributions ¶ Suppose that ( ξ , η ) (\xi,\eta) ( ξ , η ) is a pair of continuous random variables.
A joint probability density function is a nonnegative function
p ξ , η ( x , y ) p_{\xi,\eta}(x,y) p ξ , η ( x , y ) satisfying
∫ − ∞ + ∞ ∫ − ∞ + ∞ p ξ , η ( x , y ) d x d y = 1. \int_{-\infty}^{+\infty}
\int_{-\infty}^{+\infty}
p_{\xi,\eta}(x,y)\,dx\,dy
=
1. ∫ − ∞ + ∞ ∫ − ∞ + ∞ p ξ , η ( x , y ) d x d y = 1. The probability that ( ξ , η ) (\xi,\eta) ( ξ , η ) belongs to a region D ⊆ R 2 D\subseteq\mathbb{R}^2 D ⊆ R 2 is
P ( { ( ξ , η ) ∈ D } ) = ∬ D p ξ , η ( x , y ) d x d y . \mathbb{P}
\left(
\{(\xi,\eta)\in D\}
\right)
=
\iint_D
p_{\xi,\eta}(x,y)\,dx\,dy. P ( {( ξ , η ) ∈ D } ) = ∬ D p ξ , η ( x , y ) d x d y . For a rectangle,
P ( { a ≤ ξ ≤ b , c ≤ η ≤ d } ) = ∫ a b ∫ c d p ξ , η ( x , y ) d y d x . \begin{aligned}
&\mathbb{P}
\left(
\{a\leq\xi\leq b,\ c\leq\eta\leq d\}
\right)\\
&\qquad=
\int_a^b\int_c^d
p_{\xi,\eta}(x,y)\,dy\,dx.
\end{aligned} P ( { a ≤ ξ ≤ b , c ≤ η ≤ d } ) = ∫ a b ∫ c d p ξ , η ( x , y ) d y d x . Thus, just as in one dimension, probabilities correspond to areas under a density. In two dimensions, the density is a surface and probabilities correspond to volumes under that surface.
Marginal densities ¶ The marginal density of ξ \xi ξ is obtained by integrating out η \eta η :
p ξ ( x ) = ∫ − ∞ + ∞ p ξ , η ( x , y ) d y . p_\xi(x)
=
\int_{-\infty}^{+\infty}
p_{\xi,\eta}(x,y)\,dy. p ξ ( x ) = ∫ − ∞ + ∞ p ξ , η ( x , y ) d y . Similarly,
p η ( y ) = ∫ − ∞ + ∞ p ξ , η ( x , y ) d x . p_\eta(y)
=
\int_{-\infty}^{+\infty}
p_{\xi,\eta}(x,y)\,dx. p η ( y ) = ∫ − ∞ + ∞ p ξ , η ( x , y ) d x . This is the continuous analogue of summing the rows and columns of a discrete joint probability table.
The terminology “marginal” is therefore used in both cases:
P ξ ( x ) = ∑ y P ξ , η ( x , y ) , p ξ ( x ) = ∫ p ξ , η ( x , y ) d y , P η ( y ) = ∑ x P ξ , η ( x , y ) , p η ( y ) = ∫ p ξ , η ( x , y ) d x . \boxed{
\begin{aligned}
P_\xi(x)
&=
\sum_y P_{\xi,\eta}(x,y),
&
p_\xi(x)
&=
\int p_{\xi,\eta}(x,y)\,dy,
\\[6pt]
P_\eta(y)
&=
\sum_x P_{\xi,\eta}(x,y),
&
p_\eta(y)
&=
\int p_{\xi,\eta}(x,y)\,dx.
\end{aligned}} P ξ ( x ) P η ( y ) = y ∑ P ξ , η ( x , y ) , = x ∑ P ξ , η ( x , y ) , p ξ ( x ) p η ( y ) = ∫ p ξ , η ( x , y ) d y , = ∫ p ξ , η ( x , y ) d x . Let ξ \xi ξ and η \eta η be continuous random variables that are not independent .
Suppose ξ \xi ξ is uniformly distributed over the interval [ 0 , 10 ] [0, 10] [ 0 , 10 ] :
ξ ∼ U n i f o r m ( 0 , 10 ) , p ξ ( x ) = { 1 10 , 0 ≤ x ≤ 10 , 0 , otherwise. \xi \sim \mathrm{Uniform}(0, 10), \quad p_\xi(x) = \begin{cases} \frac{1}{10}, & 0 \le x \le 10, \\ 0, & \text{otherwise.} \end{cases} ξ ∼ Uniform ( 0 , 10 ) , p ξ ( x ) = { 10 1 , 0 , 0 ≤ x ≤ 10 , otherwise. Instead of η \eta η being purely independent, let its conditional distribution given ξ = x \xi = x ξ = x depend directly on x x x . Specifically, let the mean of η \eta η increase linearly with x x x :
η ∣ ξ = x ∼ N ( μ ( x ) , σ 2 ) , where μ ( x ) = x and σ = 1. \eta \mid \xi = x \;\sim\; \mathcal{N}(\mu(x), \sigma^2), \quad \text{where } \mu(x) = x \text{ and } \sigma = 1. η ∣ ξ = x ∼ N ( μ ( x ) , σ 2 ) , where μ ( x ) = x and σ = 1. The conditional probability density function is given by:
p η ∣ ξ ( y ∣ x ) = 1 2 π exp ( − ( y − x ) 2 2 ) . p_{\eta \mid \xi}(y \mid x) = \frac{1}{\sqrt{2\pi}} \exp\left(-\frac{(y - x)^2}{2}\right). p η ∣ ξ ( y ∣ x ) = 2 π 1 exp ( − 2 ( y − x ) 2 ) . By the product rule for joint densities, p ξ , η ( x , y ) = p ξ ( x ) ⋅ p η ∣ ξ ( y ∣ x ) p_{\xi,\eta}(x,y) = p_\xi(x) \cdot p_{\eta \mid \xi}(y \mid x) p ξ , η ( x , y ) = p ξ ( x ) ⋅ p η ∣ ξ ( y ∣ x ) . Thus, the joint PDF is:
p ξ , η ( x , y ) = { 1 10 2 π exp ( − ( y − x ) 2 2 ) , 0 ≤ x ≤ 10 , y ∈ R , 0 , otherwise. p_{\xi,\eta}(x,y) = \begin{cases}
\frac{1}{10 \sqrt{2\pi}} \exp\left(-\frac{(y - x)^2}{2}\right), & 0 \le x \le 10, \, y \in \mathbb{R}, \\[8pt]
0, & \text{otherwise.}
\end{cases} p ξ , η ( x , y ) = ⎩ ⎨ ⎧ 10 2 π 1 exp ( − 2 ( y − x ) 2 ) , 0 , 0 ≤ x ≤ 10 , y ∈ R , otherwise. Marginal Density p ξ ( x ) p_\xi(x) p ξ ( x ) :
Integrating the joint PDF over all y ∈ R y \in \mathbb{R} y ∈ R for x ∈ [ 0 , 10 ] x \in [0, 10] x ∈ [ 0 , 10 ] :
p ξ ( x ) = ∫ − ∞ + ∞ p ξ , η ( x , y ) d y = = 1 10 ∫ − ∞ + ∞ 1 2 π exp ( − ( y − x ) 2 2 ) d y = = 1 10 ⋅ 1 = 1 10 . \begin{split}
p_\xi(x) = & \int_{-\infty}^{+\infty} p_{\xi,\eta}(x,y) \, dy = \\ = & \frac{1}{10} \int_{-\infty}^{+\infty} \frac{1}{\sqrt{2\pi}} \exp\left(-\frac{(y - x)^2}{2}\right) dy = \\ = & \frac{1}{10} \cdot 1 = \frac{1}{10}.
\end{split} p ξ ( x ) = = = ∫ − ∞ + ∞ p ξ , η ( x , y ) d y = 10 1 ∫ − ∞ + ∞ 2 π 1 exp ( − 2 ( y − x ) 2 ) d y = 10 1 ⋅ 1 = 10 1 . Marginal Density p η ( y ) p_\eta(y) p η ( y ) :**
Integrating the joint PDF over all x ∈ [ 0 , 10 ] x \in [0, 10] x ∈ [ 0 , 10 ] :
p η ( y ) = ∫ − ∞ + ∞ p ξ , η ( x , y ) d x = = 1 10 ∫ 0 10 1 2 π exp ( − ( y − x ) 2 2 ) d x = = 1 10 [ Φ ( y ) − Φ ( y − 10 ) ] , \begin{split}
p_\eta(y) = & \int_{-\infty}^{+\infty} p_{\xi,\eta}(x,y) \, dx = \\ = & \frac{1}{10} \int_{0}^{10} \frac{1}{\sqrt{2\pi}} \exp\left(-\frac{(y - x)^2}{2}\right) dx = \\ = & \frac{1}{10} \left[ \Phi(y) - \Phi(y - 10) \right],
\end{split} p η ( y ) = = = ∫ − ∞ + ∞ p ξ , η ( x , y ) d x = 10 1 ∫ 0 10 2 π 1 exp ( − 2 ( y − x ) 2 ) d x = 10 1 [ Φ ( y ) − Φ ( y − 10 ) ] , where Φ ( ⋅ ) \Phi(\cdot) Φ ( ⋅ ) is the cumulative distribution function (CDF) of the standard normal distribution. Because p ξ , η ( x , y ) ≠ p ξ ( x ) p η ( y ) p_{\xi,\eta}(x,y) \neq p_\xi(x) p_\eta(y) p ξ , η ( x , y ) = p ξ ( x ) p η ( y ) , the random variables ξ \xi ξ and η \eta η are dependent .
import matplotlib.pyplot as plt
import numpy as np
from scipy.integrate import simpson
from scipy.stats import norm, uniform
# 1. Define distribution parameters
a, b = 0.0, 10.0 # Uniform bounds for xi
sigma = 1.0 # Standard deviation for conditional normal eta | xi = x
# 2. Create spatial grid
x = np.linspace(a - 2, b + 2, 300)
y = np.linspace(a - 4 * sigma, b + 4 * sigma, 300)
X, Y = np.meshgrid(x, y)
# 3. Compute joint PDF values p(x, y) = p_xi(x) * p_(eta|xi)(y|x)
p_xi_grid = uniform.pdf(X, loc=a, scale=b - a)
p_eta_given_xi_grid = norm.pdf(Y, loc=X, scale=sigma)
Z = p_xi_grid * p_eta_given_xi_grid
# 4. Compute marginal densities by numerical integration
# p_xi(x) = integral over y of p(x, y)
p_xi_marginal = simpson(Z, y, axis=0)
# p_eta(y) = integral over x of p(x, y)
p_eta_marginal = simpson(Z, x, axis=1)
# Analytical expressions for theoretical comparison
p_xi_analytical = uniform.pdf(x, loc=a, scale=b - a)
p_eta_analytical = (norm.cdf(y, loc=0, scale=1) - norm.cdf(y - 10, loc=0, scale=1)) / 10.0
# 5. Create 2x2 multi-panel plot
fig = plt.figure(figsize=(14, 10))
# --- ROW 1: JOINT DISTRIBUTION PLOTS ---
# Panel (1,1): 3D Surface Plot of Dependent Joint Density
ax1 = fig.add_subplot(2, 2, 1, projection='3d')
surf = ax1.plot_surface(
X, Y, Z, cmap='viridis', edgecolor='none', alpha=0.85
)
ax1.set_title(r'Joint Density Surface $p_{\xi,\eta}(x,y)$', fontsize=12, pad=12)
ax1.set_xlabel(r'$\xi$ (Uniform)')
ax1.set_ylabel(r'$\eta$ (Conditionally Normal)')
ax1.set_zlabel(r'Density $p(x,y)$')
ax1.view_init(elev=28, azim=-55)
# Panel (1,2): 2D Contour Plot
ax2 = fig.add_subplot(2, 2, 2)
contour = ax2.contourf(X, Y, Z, levels=15, cmap='viridis')
fig.colorbar(contour, ax=ax2, label=r'Density $p_{\xi,\eta}(x,y)$')
x_line = np.linspace(a, b, 100)
ax2.plot(
x_line,
x_line,
color='red',
linestyle='--',
linewidth=2,
label=r'$\mathbb{E}[\eta \mid \xi = x] = x$',
)
ax2.set_title(r'Joint Contours with Mean Trend $\mu(x) = x$', fontsize=12)
ax2.set_xlabel(r'$x$')
ax2.set_ylabel(r'$y$')
ax2.grid(True, linestyle='--', alpha=0.5)
ax2.legend(loc='upper left')
# --- ROW 2: MARGINAL DISTRIBUTION PLOTS ---
# Panel (2,1): Marginal Density p_xi(x)
ax3 = fig.add_subplot(2, 2, 3)
ax3.plot(
x,
p_xi_marginal,
color='navy',
lw=2,
label=r'Integrated $\int p(x,y)dy$',
)
ax3.plot(
x,
p_xi_analytical,
color='deepskyblue',
linestyle='--',
lw=2,
label=r'Theoretical $\mathrm{Uniform}(0,10)$',
)
ax3.fill_between(x, p_xi_marginal, color='skyblue', alpha=0.3)
ax3.set_title(r'Marginal Density $p_\xi(x)$', fontsize=12)
ax3.set_xlabel(r'$x$')
ax3.set_ylabel(r'Density $p_\xi(x)$')
ax3.grid(True, linestyle='--', alpha=0.5)
ax3.legend()
# Panel (2,2): Marginal Density p_eta(y)
ax4 = fig.add_subplot(2, 2, 4)
ax4.plot(
y,
p_eta_marginal,
color='crimson',
lw=2,
label=r'Integrated $\int p(x,y)dx$',
)
ax4.plot(
y,
p_eta_analytical,
color='orange',
linestyle='--',
lw=2,
label=r'Theoretical $\frac{1}{10}[\Phi(y)-\Phi(y-10)]$',
)
ax4.fill_between(y, p_eta_marginal, color='lightcoral', alpha=0.3)
ax4.set_title(r'Marginal Density $p_\eta(y)$', fontsize=12)
ax4.set_xlabel(r'$y$')
ax4.set_ylabel(r'Density $p_\eta(y)$')
ax4.grid(True, linestyle='--', alpha=0.5)
ax4.legend()
plt.tight_layout()
plt.show()Conditional Distributions ¶ The joint distribution describes how two random variables behave together, while the marginal distributions describe each variable separately. A third useful description is obtained when the value of one random variable is known.
Suppose that ξ \xi ξ and η \eta η are random variables. If we know that η = y \eta=y η = y , we may ask how the distribution of ξ \xi ξ changes. This leads to the notion of a conditional distribution .
Conditional distribution in the discrete case ¶ Suppose that ξ \xi ξ and η \eta η are discrete random variables. For a value y y y such that
P η ( y ) > 0 , P_\eta(y)>0, P η ( y ) > 0 , the conditional probability that ξ = x \xi=x ξ = x given η = y \eta=y η = y is
P ξ ∣ η ( x ∣ y ) = P ( { ξ = x } ∣ { η = y } ) = P ξ , η ( x , y ) P η ( y ) . P_{\xi\mid\eta}(x\mid y)
=
\mathbb{P}(\{\xi=x\}\mid\{\eta=y\})
=
\frac{P_{\xi,\eta}(x,y)}{P_\eta(y)}. P ξ ∣ η ( x ∣ y ) = P ({ ξ = x } ∣ { η = y }) = P η ( y ) P ξ , η ( x , y ) . For every fixed y y y with P η ( y ) > 0 P_\eta(y)>0 P η ( y ) > 0 , this is a probability distribution in x x x . In particular,
P ξ ∣ η ( x ∣ y ) ≥ 0 , P_{\xi\mid\eta}(x\mid y)\geq 0, P ξ ∣ η ( x ∣ y ) ≥ 0 , and
∑ x P ξ ∣ η ( x ∣ y ) = 1. \sum_x P_{\xi\mid\eta}(x\mid y)=1. x ∑ P ξ ∣ η ( x ∣ y ) = 1. The joint distribution can therefore be reconstructed from the marginal and conditional distributions:
P ξ , η ( x , y ) = P ξ ∣ η ( x ∣ y ) P η ( y ) . P_{\xi,\eta}(x,y)
=
P_{\xi\mid\eta}(x\mid y)P_\eta(y). P ξ , η ( x , y ) = P ξ ∣ η ( x ∣ y ) P η ( y ) . This is the random-variable version of the multiplication rule introduced for events in Lecture 2.
Suppose that a production line produces components classified according to their type and whether they are defective. Let
ξ = { 1 , the component is defective , 0 , otherwise , \xi=
\begin{cases}
1,&\text{the component is defective},\\
0,&\text{otherwise},
\end{cases} ξ = { 1 , 0 , the component is defective , otherwise , and let η \eta η denote the production line, with η ∈ { 1 , 2 } \eta\in\{1,2\} η ∈ { 1 , 2 } .
Suppose that the joint distribution is
η = 1 η = 2 ξ = 0 0.45 0.40 ξ = 1 0.05 0.10 \begin{array}{c|cc}
& \eta=1 & \eta=2\\
\hline
\xi=0 & 0.45 & 0.40\\
\xi=1 & 0.05 & 0.10
\end{array} ξ = 0 ξ = 1 η = 1 0.45 0.05 η = 2 0.40 0.10 Then
P η ( 1 ) = 0.50 , P η ( 2 ) = 0.50. P_\eta(1)=0.50,
\qquad
P_\eta(2)=0.50. P η ( 1 ) = 0.50 , P η ( 2 ) = 0.50. If we are told that the component came from line 1, its conditional probability of being defective is
P ξ ∣ η ( 1 ∣ 1 ) = 0.05 0.50 = 0.10. P_{\xi\mid\eta}(1\mid 1)
=
\frac{0.05}{0.50}
=
0.10. P ξ ∣ η ( 1 ∣ 1 ) = 0.50 0.05 = 0.10. Thus, conditioning on the production line changes the distribution of ξ \xi ξ .
Conditional density in the continuous case ¶ Suppose now that the pair ( ξ , η ) (\xi,\eta) ( ξ , η ) admits a joint density p ξ , η p_{\xi,\eta} p ξ , η , and let p η p_\eta p η be the marginal density of η \eta η .
For values y y y such that p η ( y ) > 0 p_\eta(y)>0 p η ( y ) > 0 , the conditional density of ξ \xi ξ given η = y \eta=y η = y is defined by
p ξ ∣ η ( x ∣ y ) = p ξ , η ( x , y ) p η ( y ) . p_{\xi\mid\eta}(x\mid y)
=
\frac{p_{\xi,\eta}(x,y)}{p_\eta(y)}. p ξ ∣ η ( x ∣ y ) = p η ( y ) p ξ , η ( x , y ) . For each fixed y y y for which the conditional density is defined,
∫ − ∞ + ∞ p ξ ∣ η ( x ∣ y ) d x = 1. \int_{-\infty}^{+\infty}
p_{\xi\mid\eta}(x\mid y)\,dx
=
1. ∫ − ∞ + ∞ p ξ ∣ η ( x ∣ y ) d x = 1. The corresponding conditional distribution function is
F ξ ∣ η ( x ∣ y ) = ∫ − ∞ x p ξ ∣ η ( t ∣ y ) d t . F_{\xi\mid\eta}(x\mid y)
=
\int_{-\infty}^{x}
p_{\xi\mid\eta}(t\mid y)\,dt. F ξ ∣ η ( x ∣ y ) = ∫ − ∞ x p ξ ∣ η ( t ∣ y ) d t . As in the discrete case, the joint density can be written as
p ξ , η ( x , y ) = p ξ ∣ η ( x ∣ y ) p η ( y ) . p_{\xi,\eta}(x,y)
=
p_{\xi\mid\eta}(x\mid y)p_\eta(y). p ξ , η ( x , y ) = p ξ ∣ η ( x ∣ y ) p η ( y ) . Although the notation p ξ ∣ η ( x ∣ y ) p_{\xi\mid\eta}(x\mid y) p ξ ∣ η ( x ∣ y ) resembles an ordinary probability, it is a density , not the probability that ξ = x \xi=x ξ = x . For a continuous random variable,
P ( { ξ = x } ) = 0 \mathbb{P}(\{\xi=x\})=0 P ({ ξ = x }) = 0 for every fixed x x x . Probabilities are obtained by integrating the conditional density over intervals or more general sets.
Independence of Random Variables ¶ For events, independence means that knowing whether one event occurred does not change the probability of another event. The same idea extends naturally to random variables.
Two random variables ξ \xi ξ and η \eta η are said to be independent if, for every pair of Borel sets A , B ⊆ R A,B\subseteq\mathbb{R} A , B ⊆ R ,
P ( { ξ ∈ A } ∩ { η ∈ B } ) = P ( { ξ ∈ A } ) P ( { η ∈ B } ) . \mathbb{P}(\{\xi\in A\}\cap\{\eta\in B\})
=
\mathbb{P}(\{\xi\in A\})
\mathbb{P}(\{\eta\in B\}). P ({ ξ ∈ A } ∩ { η ∈ B }) = P ({ ξ ∈ A }) P ({ η ∈ B }) . In words, observing the value of ξ \xi ξ provides no probabilistic information about the value of η \eta η , and vice versa.
Independence in the discrete case ¶ For discrete random variables, independence is equivalent to the factorization
P ξ , η ( x , y ) = P ξ ( x ) P η ( y ) P_{\xi,\eta}(x,y)
=
P_\xi(x)P_\eta(y) P ξ , η ( x , y ) = P ξ ( x ) P η ( y ) for every pair ( x , y ) (x,y) ( x , y ) .
For example, if ξ \xi ξ and η \eta η are independent Bernoulli random variables with parameters p p p and q q q , respectively, then
P ( { ξ = 1 , η = 1 } ) = p q , P ( { ξ = 1 , η = 0 } ) = p ( 1 − q ) , P ( { ξ = 0 , η = 1 } ) = ( 1 − p ) q , P ( { ξ = 0 , η = 0 } ) = ( 1 − p ) ( 1 − q ) . \begin{aligned}
\mathbb{P}(\{\xi=1,\eta=1\}) &= pq,\\
\mathbb{P}(\{\xi=1,\eta=0\}) &= p(1-q),\\
\mathbb{P}(\{\xi=0,\eta=1\}) &= (1-p)q,\\
\mathbb{P}(\{\xi=0,\eta=0\}) &= (1-p)(1-q).
\end{aligned} P ({ ξ = 1 , η = 1 }) P ({ ξ = 1 , η = 0 }) P ({ ξ = 0 , η = 1 }) P ({ ξ = 0 , η = 0 }) = pq , = p ( 1 − q ) , = ( 1 − p ) q , = ( 1 − p ) ( 1 − q ) . The joint distribution is completely determined by the two marginal distributions.
Independence in the continuous case ¶ If ( ξ , η ) (\xi,\eta) ( ξ , η ) admits a joint density, independence is equivalent to
p ξ , η ( x , y ) = p ξ ( x ) p η ( y ) p_{\xi,\eta}(x,y)
=
p_\xi(x)p_\eta(y) p ξ , η ( x , y ) = p ξ ( x ) p η ( y ) for almost every ( x , y ) ∈ R 2 (x,y)\in\mathbb{R}^2 ( x , y ) ∈ R 2 .
Equivalently, the joint distribution function factorizes as
F ξ , η ( x , y ) = F ξ ( x ) F η ( y ) . F_{\xi,\eta}(x,y)
=
F_\xi(x)F_\eta(y). F ξ , η ( x , y ) = F ξ ( x ) F η ( y ) . Thus, independence is a special situation in which the joint distribution contains no additional information beyond the two marginal distributions.
Independence and conditional distributions ¶ The connection with conditional distributions is particularly useful.
If ξ \xi ξ and η \eta η are independent, then, whenever the conditional distribution is defined,
P ξ ∣ η ( x ∣ y ) = P ξ ( x ) P_{\xi\mid\eta}(x\mid y)
=
P_\xi(x) P ξ ∣ η ( x ∣ y ) = P ξ ( x ) in the discrete case, and
p ξ ∣ η ( x ∣ y ) = p ξ ( x ) p_{\xi\mid\eta}(x\mid y)
=
p_\xi(x) p ξ ∣ η ( x ∣ y ) = p ξ ( x ) in the continuous case.
Thus, conditioning on η \eta η does not change the distribution of ξ \xi ξ .
The converse also holds under the usual conditions: if the conditional distribution of ξ \xi ξ does not depend on y y y , then ξ \xi ξ and η \eta η are independent.
Functions of Two Random Variables ¶ Once a joint distribution is available, we can compute expectations not only of ξ \xi ξ and η \eta η separately, but also of functions involving both variables.
Let
g : R 2 → R g:\mathbb{R}^2\to\mathbb{R} g : R 2 → R be a suitable function.
If ξ \xi ξ and η \eta η are discrete, then
E [ g ( ξ , η ) ] = ∑ x ∑ y g ( x , y ) P ξ , η ( x , y ) . \mathbb{E}[g(\xi,\eta)]
=
\sum_x\sum_y
g(x,y)P_{\xi,\eta}(x,y). E [ g ( ξ , η )] = x ∑ y ∑ g ( x , y ) P ξ , η ( x , y ) . If ( ξ , η ) (\xi,\eta) ( ξ , η ) admits a joint density, then
E [ g ( ξ , η ) ] = ∫ − ∞ + ∞ ∫ − ∞ + ∞ g ( x , y ) p ξ , η ( x , y ) d x d y . \mathbb{E}[g(\xi,\eta)]
=
\int_{-\infty}^{+\infty}
\int_{-\infty}^{+\infty}
g(x,y)p_{\xi,\eta}(x,y)\,dx\,dy. E [ g ( ξ , η )] = ∫ − ∞ + ∞ ∫ − ∞ + ∞ g ( x , y ) p ξ , η ( x , y ) d x d y . These formulas are direct extensions of the one-dimensional formulas introduced in Lecture 6.
For example, choosing
g ( x , y ) = x y g(x,y)=xy g ( x , y ) = x y gives the mixed first moment
E [ ξ η ] . \mathbb{E}[\xi\eta]. E [ ξ η ] . More generally, the quantities
E [ ξ r η s ] , r , s ∈ N , \mathbb{E}[\xi^r\eta^s],
\qquad r,s\in\mathbb{N}, E [ ξ r η s ] , r , s ∈ N , are called joint moments of ξ \xi ξ and η \eta η .
The exponents determine the order of the moment. For example,
E [ ξ η ] \mathbb{E}[\xi\eta] E [ ξ η ] is a mixed moment of total order two, while
E [ ξ 2 η ] \mathbb{E}[\xi^2\eta] E [ ξ 2 η ] has total order three.
Products of independent random variables ¶ Independence leads to an important simplification.
If ξ \xi ξ and η \eta η are independent and the expectations exist, then
E [ ξ η ] = E [ ξ ] E [ η ] . \mathbb{E}[\xi\eta]
=
\mathbb{E}[\xi]\mathbb{E}[\eta]. E [ ξ η ] = E [ ξ ] E [ η ] . More generally, if the relevant expectations exist,
E [ g ( ξ ) h ( η ) ] = E [ g ( ξ ) ] E [ h ( η ) ] . \mathbb{E}[g(\xi)h(\eta)]
=
\mathbb{E}[g(\xi)]\mathbb{E}[h(\eta)]. E [ g ( ξ ) h ( η )] = E [ g ( ξ )] E [ h ( η )] . For the discrete case, this follows immediately from the factorization of the joint pmf:
E [ g ( ξ ) h ( η ) ] = ∑ x ∑ y g ( x ) h ( y ) P ξ , η ( x , y ) = ∑ x ∑ y g ( x ) h ( y ) P ξ ( x ) P η ( y ) = ( ∑ x g ( x ) P ξ ( x ) ) ( ∑ y h ( y ) P η ( y ) ) . \begin{aligned}
\mathbb{E}[g(\xi)h(\eta)]
&=
\sum_x\sum_y
g(x)h(y)P_{\xi,\eta}(x,y)\\
&=
\sum_x\sum_y
g(x)h(y)P_\xi(x)P_\eta(y)\\
&=
\left(\sum_x g(x)P_\xi(x)\right)
\left(\sum_y h(y)P_\eta(y)\right).
\end{aligned} E [ g ( ξ ) h ( η )] = x ∑ y ∑ g ( x ) h ( y ) P ξ , η ( x , y ) = x ∑ y ∑ g ( x ) h ( y ) P ξ ( x ) P η ( y ) = ( x ∑ g ( x ) P ξ ( x ) ) ( y ∑ h ( y ) P η ( y ) ) . The continuous case follows in the same way by using the factorization of the joint density.
Covariance ¶ The product E [ ξ η ] \mathbb{E}[\xi\eta] E [ ξ η ] alone does not tell us whether large values of ξ \xi ξ tend to occur together with large values of η \eta η . To measure this type of joint variation, we introduce the covariance .
Assume that ξ \xi ξ and η \eta η have finite second moments. Their covariance is defined by
Cov ( ξ , η ) = E [ ( ξ − E [ ξ ] ) ( η − E [ η ] ) ] . \operatorname{Cov}(\xi,\eta)
=
\mathbb{E}
\left[
(\xi-\mathbb{E}[\xi])
(\eta-\mathbb{E}[\eta])
\right]. Cov ( ξ , η ) = E [ ( ξ − E [ ξ ]) ( η − E [ η ]) ] . The covariance measures whether the two variables tend to deviate from their means in the same direction.
Expanding the product gives
Cov ( ξ , η ) = E [ ξ η ] − E [ ξ ] E [ η ] . \operatorname{Cov}(\xi,\eta)
=
\mathbb{E}[\xi\eta]
-
\mathbb{E}[\xi]\mathbb{E}[\eta]. Cov ( ξ , η ) = E [ ξ η ] − E [ ξ ] E [ η ] . Thus, covariance compares the mixed moment E [ ξ η ] \mathbb{E}[\xi\eta] E [ ξ η ] with the product of the means.
If large values of ξ \xi ξ tend to be associated with large values of η \eta η , the covariance is typically positive. If large values of one variable tend to be associated with small values of the other, it is typically negative.
The covariance is symmetric:
Cov ( ξ , η ) = Cov ( η , ξ ) . \operatorname{Cov}(\xi,\eta)
=
\operatorname{Cov}(\eta,\xi). Cov ( ξ , η ) = Cov ( η , ξ ) . It is also linear in each argument. For constants a , b , c , d a,b,c,d a , b , c , d ,
Cov ( a ξ + b , c η + d ) = a c Cov ( ξ , η ) . \operatorname{Cov}(a\xi+b,c\eta+d)
=
ac\,\operatorname{Cov}(\xi,\eta). Cov ( a ξ + b , cη + d ) = a c Cov ( ξ , η ) . In particular,
Cov ( ξ , c ) = 0 \operatorname{Cov}(\xi,c)=0 Cov ( ξ , c ) = 0 for every constant c c c .
Most importantly, variance is a special case of covariance:
Cov ( ξ , ξ ) = Var ( ξ ) . \operatorname{Cov}(\xi,\xi)
=
\operatorname{Var}(\xi). Cov ( ξ , ξ ) = Var ( ξ ) . Variance of a sum ¶ Using covariance, the variance of a sum can be written as
Var ( ξ + η ) = Var ( ξ ) + Var ( η ) + 2 Cov ( ξ , η ) . \operatorname{Var}(\xi+\eta)
=
\operatorname{Var}(\xi)
+
\operatorname{Var}(\eta)
+
2\operatorname{Cov}(\xi,\eta). Var ( ξ + η ) = Var ( ξ ) + Var ( η ) + 2 Cov ( ξ , η ) . More generally,
Var ( ∑ k = 1 n ξ k ) = ∑ k = 1 n Var ( ξ k ) + 2 ∑ 1 ≤ i < j ≤ n Cov ( ξ i , ξ j ) . \operatorname{Var}
\left(
\sum_{k=1}^n \xi_k
\right)
=
\sum_{k=1}^n\operatorname{Var}(\xi_k)
+
2\sum_{1\leq i<j\leq n}
\operatorname{Cov}(\xi_i,\xi_j). Var ( k = 1 ∑ n ξ k ) = k = 1 ∑ n Var ( ξ k ) + 2 1 ≤ i < j ≤ n ∑ Cov ( ξ i , ξ j ) . If the random variables are pairwise uncorrelated, all covariance terms vanish and therefore
Var ( ∑ k = 1 n ξ k ) = ∑ k = 1 n Var ( ξ k ) . \operatorname{Var}
\left(
\sum_{k=1}^n \xi_k
\right)
=
\sum_{k=1}^n\operatorname{Var}(\xi_k). Var ( k = 1 ∑ n ξ k ) = k = 1 ∑ n Var ( ξ k ) . In particular, independent random variables are uncorrelated whenever their second moments exist. Hence the variance formula for sums of independent random variables introduced in Lecture 6 is a consequence of the more general covariance formula.
Correlation ¶ The numerical value of covariance depends on the units in which the variables are measured. For example, changing a measurement from metres to centimetres multiplies the covariance by a factor of 102 .
To obtain a dimensionless measure, we normalize the covariance by the standard deviations.
Suppose that
0 < Var ( ξ ) < ∞ , 0 < Var ( η ) < ∞ . 0<\operatorname{Var}(\xi)<\infty,
\qquad
0<\operatorname{Var}(\eta)<\infty. 0 < Var ( ξ ) < ∞ , 0 < Var ( η ) < ∞. The correlation coefficient of ξ \xi ξ and η \eta η is
ρ ξ , η = Cov ( ξ , η ) Var ( ξ ) Var ( η ) . \rho_{\xi,\eta}
=
\frac{\operatorname{Cov}(\xi,\eta)}
{\sqrt{\operatorname{Var}(\xi)}
\sqrt{\operatorname{Var}(\eta)}}. ρ ξ , η = Var ( ξ ) Var ( η ) Cov ( ξ , η ) . Using the standard deviations σ ξ \sigma_\xi σ ξ and σ η \sigma_\eta σ η , this can be written as
ρ ξ , η = Cov ( ξ , η ) σ ξ σ η . \rho_{\xi,\eta}
=
\frac{\operatorname{Cov}(\xi,\eta)}
{\sigma_\xi\sigma_\eta}. ρ ξ , η = σ ξ σ η Cov ( ξ , η ) . The Cauchy--Schwarz inequality implies that
− 1 ≤ ρ ξ , η ≤ 1. -1\leq \rho_{\xi,\eta}\leq 1. − 1 ≤ ρ ξ , η ≤ 1. The value of ρ ξ , η \rho_{\xi,\eta} ρ ξ , η describes the strength and direction of a linear association between the two variables:
ρ ξ , η > 0 \rho_{\xi,\eta}>0 ρ ξ , η > 0 indicates positive linear association;
ρ ξ , η < 0 \rho_{\xi,\eta}<0 ρ ξ , η < 0 indicates negative linear association;
ρ ξ , η = 0 \rho_{\xi,\eta}=0 ρ ξ , η = 0 means that the variables are uncorrelated;
values close to 1 or -1 indicate a strong linear relationship.
It is important to distinguish correlation from independence.
ξ and η independent ⟹ ρ ξ , η = 0 , \xi\text{ and }\eta\text{ independent}
\quad\Longrightarrow\quad
\rho_{\xi,\eta}=0, ξ and η independent ⟹ ρ ξ , η = 0 , provided the variances are finite and positive.
The converse is generally false.
Zero correlation does not imply independence ¶ Consider a random variable
ξ ∼ U ( − 1 , 1 ) \xi\sim U(-1,1) ξ ∼ U ( − 1 , 1 ) and define
Clearly, η \eta η is completely determined by ξ \xi ξ , so the two variables cannot be independent.
However, by symmetry,
E [ ξ ] = 0 \mathbb{E}[\xi]=0 E [ ξ ] = 0 and
E [ ξ 3 ] = 0. \mathbb{E}[\xi^3]=0. E [ ξ 3 ] = 0. Consequently,
Cov ( ξ , η ) = Cov ( ξ , ξ 2 ) = E [ ξ 3 ] − E [ ξ ] E [ ξ 2 ] = 0. \operatorname{Cov}(\xi,\eta)
=
\operatorname{Cov}(\xi,\xi^2)
=
\mathbb{E}[\xi^3]
-
\mathbb{E}[\xi]\mathbb{E}[\xi^2]
=
0. Cov ( ξ , η ) = Cov ( ξ , ξ 2 ) = E [ ξ 3 ] − E [ ξ ] E [ ξ 2 ] = 0. Thus ξ \xi ξ and η \eta η are uncorrelated but not independent.
This example illustrates an important distinction:
independence ⟹ uncorrelatedness , \text{independence}
\quad\Longrightarrow\quad
\text{uncorrelatedness}, independence ⟹ uncorrelatedness , but, in general,
uncorrelatedness ⟹̸ independence . \text{uncorrelatedness}
\quad\not\Longrightarrow\quad
\text{independence}. uncorrelatedness ⟹ independence . Correlation detects only a particular type of dependence, namely linear dependence.
An Interpretation Through Linear Prediction ¶ The correlation coefficient also has a useful interpretation in terms of linear prediction.
Suppose we want to approximate ξ \xi ξ by an affine function of η \eta η ,
ξ ^ = c 1 + c 2 η . \widehat{\xi}=c_1+c_2\eta. ξ = c 1 + c 2 η . We choose c 1 c_1 c 1 and c 2 c_2 c 2 to minimize the mean squared error
E [ ( ξ − c 1 − c 2 η ) 2 ] . \mathbb{E}
\left[
(\xi-c_1-c_2\eta)^2
\right]. E [ ( ξ − c 1 − c 2 η ) 2 ] . Assume that ξ \xi ξ and η \eta η have finite, positive variances. The optimal coefficients are
c 2 = Cov ( ξ , η ) Var ( η ) c_2
=
\frac{\operatorname{Cov}(\xi,\eta)}
{\operatorname{Var}(\eta)} c 2 = Var ( η ) Cov ( ξ , η ) and
c 1 = E [ ξ ] − c 2 E [ η ] . c_1
=
\mathbb{E}[\xi]
-
c_2\mathbb{E}[\eta]. c 1 = E [ ξ ] − c 2 E [ η ] . The corresponding minimum mean squared error is
min c 1 , c 2 E [ ( ξ − c 1 − c 2 η ) 2 ] = Var ( ξ ) ( 1 − ρ ξ , η 2 ) . \min_{c_1,c_2}
\mathbb{E}
\left[
(\xi-c_1-c_2\eta)^2
\right]
=
\operatorname{Var}(\xi)(1-\rho_{\xi,\eta}^2). c 1 , c 2 min E [ ( ξ − c 1 − c 2 η ) 2 ] = Var ( ξ ) ( 1 − ρ ξ , η 2 ) . Thus ρ ξ , η 2 \rho_{\xi,\eta}^2 ρ ξ , η 2 measures the fraction of the variance of ξ \xi ξ that can be explained by the best affine prediction based on η \eta η .
This interpretation also clarifies why zero correlation does not necessarily mean independence. If ρ ξ , η = 0 \rho_{\xi,\eta}=0 ρ ξ , η = 0 , the best affine predictor of ξ \xi ξ based on η \eta η is simply its mean. Nevertheless, nonlinear relationships between the variables may still exist, as in the example η = ξ 2 \eta=\xi^2 η = ξ 2 .
Summary ¶ The joint distribution provides a complete probabilistic description of two random variables and allows us to study their dependence.
For discrete random variables, the joint distribution is described by
P ξ , η ( x , y ) = P ( { ξ = x , η = y } ) , P_{\xi,\eta}(x,y)
=
\mathbb{P}(\{\xi=x,\eta=y\}), P ξ , η ( x , y ) = P ({ ξ = x , η = y }) , while for a pair admitting a joint density it is described by p ξ , η ( x , y ) p_{\xi,\eta}(x,y) p ξ , η ( x , y ) .
From the joint distribution we can obtain:
Object Discrete case Continuous case Joint distribution P ξ , η ( x , y ) P_{\xi,\eta}(x,y) P ξ , η ( x , y ) p ξ , η ( x , y ) p_{\xi,\eta}(x,y) p ξ , η ( x , y ) Marginal of ξ \xi ξ P ξ ( x ) = ∑ y P ξ , η ( x , y ) \displaystyle P_\xi(x)=\sum_yP_{\xi,\eta}(x,y) P ξ ( x ) = y ∑ P ξ , η ( x , y ) p ξ ( x ) = ∫ p ξ , η ( x , y ) d y \displaystyle p_\xi(x)=\int p_{\xi,\eta}(x,y)\,dy p ξ ( x ) = ∫ p ξ , η ( x , y ) d y Conditional distribution P ξ ∣ η ( x ∣ y ) = P ξ , η ( x , y ) P η ( y ) \displaystyle P_{\xi\mid\eta}(x\mid y)=\frac{P_{\xi,\eta}(x,y)}{P_\eta(y)} P ξ ∣ η ( x ∣ y ) = P η ( y ) P ξ , η ( x , y ) p ξ ∣ η ( x ∣ y ) = p ξ , η ( x , y ) p η ( y ) \displaystyle p_{\xi\mid\eta}(x\mid y)=\frac{p_{\xi,\eta}(x,y)}{p_\eta(y)} p ξ ∣ η ( x ∣ y ) = p η ( y ) p ξ , η ( x , y ) Independence P ξ , η = P ξ P η P_{\xi,\eta}=P_\xi P_\eta P ξ , η = P ξ P η p ξ , η = p ξ p η p_{\xi,\eta}=p_\xi p_\eta p ξ , η = p ξ p η Joint moment ∑ x ∑ y x r y s P ξ , η ( x , y ) \displaystyle \sum_x\sum_y x^ry^sP_{\xi,\eta}(x,y) x ∑ y ∑ x r y s P ξ , η ( x , y ) ∬ x r y s p ξ , η ( x , y ) d x d y \displaystyle \iint x^ry^s p_{\xi,\eta}(x,y)\,dx\,dy ∬ x r y s p ξ , η ( x , y ) d x d y Covariance Cov ( ξ , η ) = E [ ξ η ] − E [ ξ ] E [ η ] \operatorname{Cov}(\xi,\eta)=\mathbb{E}[\xi\eta]-\mathbb{E}[\xi]\mathbb{E}[\eta] Cov ( ξ , η ) = E [ ξ η ] − E [ ξ ] E [ η ] same Correlation ρ ξ , η = Cov ( ξ , η ) σ ξ σ η \displaystyle \rho_{\xi,\eta}=\frac{\operatorname{Cov}(\xi,\eta)}{\sigma_\xi\sigma_\eta} ρ ξ , η = σ ξ σ η Cov ( ξ , η ) same
The main relationships to remember are
joint distribution ⟶ { marginal distributions , conditional distributions , joint moments . \text{joint distribution}
\quad\longrightarrow\quad
\begin{cases}
\text{marginal distributions},\\
\text{conditional distributions},\\
\text{joint moments}.
\end{cases} joint distribution ⟶ ⎩ ⎨ ⎧ marginal distributions , conditional distributions , joint moments . Independence is characterized by factorization of the joint distribution:
P ξ , η ( x , y ) = P ξ ( x ) P η ( y ) P_{\xi,\eta}(x,y)=P_\xi(x)P_\eta(y) P ξ , η ( x , y ) = P ξ ( x ) P η ( y ) in the discrete case, or
p ξ , η ( x , y ) = p ξ ( x ) p η ( y ) p_{\xi,\eta}(x,y)=p_\xi(x)p_\eta(y) p ξ , η ( x , y ) = p ξ ( x ) p η ( y ) in the continuous case.
Finally,
independence ⟹ Cov ( ξ , η ) = 0 , \text{independence}
\quad\Longrightarrow\quad
\operatorname{Cov}(\xi,\eta)=0, independence ⟹ Cov ( ξ , η ) = 0 , but zero covariance does not generally imply independence.
From Two Random Variables to Many Observations ¶ The concepts introduced in this lecture provide the language needed to study collections of random variables.
In particular, we can now distinguish between:
the distribution of each individual observation;
the joint distribution of several observations;
independence between observations;
the moments of individual observations;
the covariance between different observations.
These ideas become especially important when considering a sequence of random variables
ξ 1 , ξ 2 , … , ξ n , … \xi_1,\xi_2,\ldots,\xi_n,\ldots ξ 1 , ξ 2 , … , ξ n , … and quantities such as their sum
S n = ∑ k = 1 n ξ k S_n=\sum_{k=1}^n\xi_k S n = k = 1 ∑ n ξ k or their average
ξ ‾ n = 1 n ∑ k = 1 n ξ k . \overline{\xi}_n
=
\frac{1}{n}\sum_{k=1}^n\xi_k. ξ n = n 1 k = 1 ∑ n ξ k . The variance calculation from Lecture 6 shows that, for independent identically distributed random variables with variance σ 2 \sigma^2 σ 2 ,
Var ( ξ ‾ n ) = σ 2 n . \operatorname{Var}(\overline{\xi}_n)
=
\frac{\sigma^2}{n}. Var ( ξ n ) = n σ 2 . Thus the average becomes increasingly concentrated around its mean as the number of observations grows.
In the next lecture we will study this phenomenon more systematically. The Law of Large Numbers describes the convergence of averages to their expected value, while the Central Limit Theorem describes the fluctuations around that limiting value and explains the fundamental role of the normal distribution in statistics and scientific computation.