Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

A probability model describes the possible outcomes of a random experiment and assigns probabilities to events. In many applications, however, we are not interested in the elementary outcomes themselves. Instead, we want to associate a numerical quantity with each outcome.

For example, when tossing a die, we may simply be interested in the number obtained. When measuring the rotational speed of a machine, the outcome of the experiment may be described by a real number. When recording the lifetime of a component, the relevant quantity is again numerical.

This motivates the concept of a random variable.

Random variables

Given a sample space Ω\Omega, a random variable is a numerical function

ξ=ξ(ω)\xi=\xi(\omega)

that assigns a real number to each elementary outcome ω∈Ω\omega\in\Omega. Thus, the value taken by ξ\xi depends on the outcome of the random experiment.

For any two real numbers x′x' and x′′x'' with x′<x′′x'<x'', consider the event

{x′<ξ≤x′′}.\{x'<\xi\leq x''\}.

This is the event consisting of all elementary outcomes ω∈Ω\omega\in\Omega for which the corresponding value ξ(ω)\xi(\omega) lies in the interval (x′,x′′](x',x'']. Its probability,

P({x′<ξ≤x′′}),\mathbb{P}(\{x'<\xi\leq x''\}),

therefore represents the probability that the random variable ξ\xi takes a value in the interval (x′,x′′](x',x''].

Knowledge of these probabilities for all intervals (x′,x′′](x',x''] completely determines the probability distribution of the random variable ξ\xi. In other words, the probability distribution specifies how the probability is distributed among the possible values of ξ\xi.

The way in which this distribution is described depends on the type of random variable. We first consider discrete random variables and then random variables having a probability density.

Discrete random variables

A random variable ξ\xi is called discrete if it takes values in a finite or countably infinite set. Denote the possible values of ξ\xi by xx.

The probability assigned to each possible value is described by the probability mass function (PMF):

Pξ(x)=P({ξ=x}).P_\xi(x)=\mathbb{P}(\{\xi=x\}).

Since ξ\xi must take one of its possible values, these probabilities satisfy

Pξ(x)≥0,∑xPξ(x)=1,P_\xi(x)\geq0, \qquad \sum_xP_\xi(x)=1,

where the sum is taken over all distinct values that ξ\xi can assume.

For a discrete random variable, the probability that ξ\xi takes a value in an interval is obtained by summing the probabilities of all possible values contained in that interval. Thus,

P({x′<ξ≤x′′})=∑x′<x≤x′′Pξ(x).\mathbb{P}(\{x'<\xi\leq x''\}) = \sum_{\substack{x'<x\leq x''}}P_\xi(x).

The PMF therefore completely determines the probability distribution of a discrete random variable.

Example: a fair die

Consider the experiment of throwing a fair six-sided die. Let ξ\xi denote the number appearing on the upper face. Then

ξ∈{1,2,3,4,5,6},\xi\in\{1,2,3,4,5,6\},

and, because the six outcomes are equally likely,

Pξ(k)=16,k=1,…,6.P_\xi(k)=\frac16, \qquad k=1,\ldots,6.

For example,

P({2≤ξ≤4})=Pξ(2)+Pξ(3)+Pξ(4)=36=12.\mathbb{P}(\{2\leq\xi\leq4\}) = P_\xi(2)+P_\xi(3)+P_\xi(4) = \frac36 = \frac12.

Thus, probabilities for intervals are obtained by summing the corresponding point probabilities.

Continuous random variables and probability densities

A random variable may instead take values throughout an interval or, more generally, throughout a continuous subset of R\mathbb{R}. In this case, it is useful to describe the distribution through a density.

More precisely, we say that a random variable ξ\xi has a probability density function (PDF), or that its distribution is absolutely continuous, if there exists a non-negative integrable function pξ:R→Rp_\xi:\mathbb{R}\to\mathbb{R} such that, for every x′<x′′x'<x'',

P({x′<ξ≤x′′})=∫x′x′′pξ(x) dx.\mathbb{P}(\{x'<\xi\leq x''\}) = \int_{x'}^{x''}p_\xi(x)\,dx.

Since the total probability must be one,

∫−∞∞pξ(x) dx=1.\int_{-\infty}^{\infty}p_\xi(x)\,dx=1.

Unlike a discrete random variable, a random variable having a density assigns zero probability to any individual value. Indeed, for every x∈Rx\in\mathbb{R},

P({ξ=x})=0.\mathbb{P}(\{\xi=x\})=0.

The density pξ(x)p_\xi(x) should therefore not be interpreted as the probability that ξ\xi takes the value xx. Rather, it describes how probability is distributed locally around xx.

If pξp_\xi is continuous at xx, then, for a small positive Δx\Delta x,

P({x<ξ≤x+Δx})=pξ(x)Δx+o(Δx).\mathbb{P}(\{x<\xi\leq x+\Delta x\}) = p_\xi(x)\Delta x+o(\Delta x).

Thus, for a sufficiently small interval, the probability is approximately the density at xx multiplied by the length of the interval.

The distribution function

The distribution function, also called the cumulative distribution function (CDF), of a random variable ξ\xi is defined by

Φξ(x)=P({ξ≤x}),−∞<x<∞.\Phi_\xi(x) = \mathbb{P}(\{\xi\leq x\}), \qquad -\infty<x<\infty.

The CDF gives the probability that the random variable does not exceed a prescribed value xx.

The distribution function is defined for every random variable, whether discrete, continuous, or of a more general type.

For every random variable, the CDF is non-decreasing and satisfies

lim⁡x→−∞Φξ(x)=0,lim⁡x→+∞Φξ(x)=1.\lim_{x\to-\infty}\Phi_\xi(x)=0, \qquad \lim_{x\to+\infty}\Phi_\xi(x)=1.

It is also right-continuous:

lim⁡h↓0Φξ(x+h)=Φξ(x).\lim_{h\downarrow0}\Phi_\xi(x+h)=\Phi_\xi(x).

Recovering probabilities from the CDF

The cumulative distribution function contains all the information needed to determine probabilities involving a random variable. By definition,

Φξ(x)=P({ξ≤x}).\Phi_\xi(x)=\mathbb{P}(\{\xi\leq x\}).

Consequently, probabilities of events defined by intervals can be obtained directly from differences of CDF values. For example,

P({ξ>x})=1−Φξ(x),\mathbb{P}(\{\xi>x\})=1-\Phi_\xi(x),

and, for x1<x2x_1<x_2,

P({x1<ξ≤x2})=Φξ(x2)−Φξ(x1).\mathbb{P}(\{x_1<\xi\leq x_2\}) = \Phi_\xi(x_2)-\Phi_\xi(x_1).

The choice of strict or non-strict inequalities at the endpoints is important for a general random variable. In particular,

P({ξ=x})=Φξ(x)−lim⁡t↑xΦξ(t).\mathbb{P}(\{\xi=x\}) = \Phi_\xi(x)-\lim_{t\uparrow x}\Phi_\xi(t).

Thus, a jump of the CDF at xx represents a positive probability concentrated at the single value xx. Such a point is called an atom of the distribution.

For example, consider a random variable ξ\xi such that

P({ξ=1})=0.2,P({ξ=2})=0.5,P({ξ=3})=0.3.\mathbb{P}(\{\xi=1\})=0.2, \qquad \mathbb{P}(\{\xi=2\})=0.5, \qquad \mathbb{P}(\{\xi=3\})=0.3.

Its CDF is a step function, and the jump at x=2x=2 has size 0.5. Hence,

P({ξ=2})=0.5.\mathbb{P}(\{\xi=2\})=0.5.

More generally, for x1<x2x_1<x_2,

P({x1≤ξ≤x2})=Φξ(x2)−lim⁡t↑x1Φξ(t),\mathbb{P}(\{x_1\leq\xi\leq x_2\}) = \Phi_\xi(x_2) - \lim_{t\uparrow x_1}\Phi_\xi(t),

while

P({x1<ξ<x2})=lim⁡t↑x2Φξ(t)−Φξ(x1).\mathbb{P}(\{x_1<\xi<x_2\}) = \lim_{t\uparrow x_2}\Phi_\xi(t) - \Phi_\xi(x_1).

For continuous random variables with a density, individual points have probability zero. Therefore, the distinction between strict and non-strict inequalities at the endpoints disappears. In that case,

P({x1≤ξ≤x2})=P({x1<ξ≤x2})=Φξ(x2)−Φξ(x1).\mathbb{P}(\{x_1\leq\xi\leq x_2\}) = \mathbb{P}(\{x_1<\xi\leq x_2\}) = \Phi_\xi(x_2)-\Phi_\xi(x_1).

This illustrates an important general principle: the CDF can be used to compute probabilities for every random variable, whereas a probability mass function is specific to discrete random variables and a probability density is available only for distributions that are absolutely continuous.

The discrete case

If ξ\xi is discrete with PMF PξP_\xi, then its CDF is obtained by summing the probabilities of all possible values that do not exceed xx:

Φξ(x)=∑t≤xPξ(t).\Phi_\xi(x) = \sum_{t\leq x}P_\xi(t).

The CDF is therefore a non-decreasing step function. It remains constant between two consecutive possible values of ξ\xi and jumps at each value that has positive probability.

The continuous case

If ξ\xi has a probability density pξp_\xi, then its CDF is obtained by integrating the density:

Φξ(x)=∫−∞xpξ(t) dt.\Phi_\xi(x) = \int_{-\infty}^{x}p_\xi(t)\,dt.

In particular, the CDF is continuous. If the density is sufficiently regular, then

pξ(x)=ddxΦξ(x).p_\xi(x)=\frac{d}{dx}\Phi_\xi(x).

Thus, the PMF, PDF, and CDF provide different ways of describing a probability distribution:

The discrete CDF is a step function because the random variable can only take isolated values. The continuous CDF is continuous because the distribution has a density.

import matplotlib.pyplot as plt
import numpy as np
from scipy.stats import binom, norm

fig, (ax1, ax2) = plt.subplots(
    1, 2, figsize=(12, 5), layout="constrained"
)

# Discrete CDF: Binomial distribution
n, p = 10, 0.5
x_discrete = np.arange(0, n + 1)
cdf_discrete = binom.cdf(x_discrete, n, p)

ax1.step(
    x_discrete,
    cdf_discrete,
    where="post",
    linewidth=2,
    label="Binomial(10, 0.5) CDF",
)
ax1.plot(x_discrete, cdf_discrete, "o", alpha=0.7)

ax1.set_title("Discrete CDF")
ax1.set_xlabel(r"$x$")
ax1.set_ylabel(r"$\Phi_\xi(x)=P(\xi\leq x)$")
ax1.set_ylim(-0.05, 1.05)
ax1.grid(True, linestyle="--", alpha=0.6)
ax1.legend(loc="lower right")

# Continuous CDF: standard normal distribution
x_continuous = np.linspace(-4, 4, 1000)
cdf_continuous = norm.cdf(x_continuous)

ax2.plot(
    x_continuous,
    cdf_continuous,
    linewidth=2,
    label=r"Standard Normal $\mathcal{N}(0,1)$ CDF",
)

ax2.set_title("Continuous CDF")
ax2.set_xlabel(r"$x$")
ax2.set_ylabel(r"$\Phi_\xi(x)=P(\xi\leq x)$")
ax2.set_ylim(-0.05, 1.05)
ax2.grid(True, linestyle="--", alpha=0.6)
ax2.legend(loc="lower right")

plt.show()
<Figure size 1200x500 with 2 Axes>

The figure illustrates the fundamental difference between the two cases. In the discrete case, the CDF changes through jumps, while in the continuous case it changes continuously.

Uniform distributions

A particularly simple probability distribution is the uniform distribution. It models a situation in which all values in a specified interval are equally likely in the sense that equal-length intervals have equal probability.

A continuous random variable XX is said to be uniformly distributed on [a,b][a,b], with a<ba<b, and we write

X∼U(a,b),X\sim U(a,b),

if its density is constant on [a,b][a,b] and zero elsewhere. Since the total area under the density must be one,

pX(x)={1b−a,a≤x≤b,0,otherwise.p_X(x) = \begin{cases} \dfrac{1}{b-a}, & a\leq x\leq b,\\[6pt] 0, & \text{otherwise}. \end{cases}

For a≤x′<x′′≤ba\leq x'<x''\leq b,

P({x′<X≤x′′})=∫x′x′′1b−a dx=x′′−x′b−a.\mathbb{P}(\{x'<X\leq x''\}) = \int_{x'}^{x''}\frac{1}{b-a}\,dx = \frac{x''-x'}{b-a}.

Thus, for a uniform distribution, the probability of an interval is proportional to its length.

The corresponding CDF is

ΦX(x)={0,x<a,x−ab−a,a≤x≤b,1,x>b.\Phi_X(x) = \begin{cases} 0, & x<a,\\[4pt] \dfrac{x-a}{b-a}, & a\leq x\leq b,\\[6pt] 1, & x>b. \end{cases}

The CDF is therefore constant before the interval [a,b][a,b], increases linearly inside the interval, and is equal to one after the interval.

A discrete analogue is obtained by assigning equal probability to a finite set of values. If XX takes the nn distinct values

x1,…,xnx_1,\ldots,x_n

with equal probability, then

P({X=xi})=1n,i=1,…,n.\mathbb{P}(\{X=x_i\})=\frac1n, \qquad i=1,\ldots,n.

This is called a discrete uniform distribution.

import matplotlib.pyplot as plt
import numpy as np
from scipy.stats import randint, uniform

fig, axs = plt.subplots(
    2, 2, figsize=(12, 8), layout="constrained"
)

# Continuous Uniform U(a,b)
a, b = 2, 8
x_cont = np.linspace(0, 10, 1000)

pdf_cont = uniform.pdf(x_cont, loc=a, scale=b - a)
cdf_cont = uniform.cdf(x_cont, loc=a, scale=b - a)

axs[0, 0].plot(
    x_cont,
    pdf_cont,
    linewidth=2,
    label=rf"PDF: $\mathcal{{U}}({a},{b})$",
)
axs[0, 0].fill_between(
    x_cont,
    pdf_cont,
    where=(x_cont >= a) & (x_cont <= b),
    alpha=0.2,
)

axs[0, 0].set_title("Continuous Uniform PDF")
axs[0, 0].set_xlabel(r"$x$")
axs[0, 0].set_ylabel(r"$p_X(x)$")
axs[0, 0].set_ylim(-0.02, 0.25)
axs[0, 0].grid(True, linestyle="--", alpha=0.6)
axs[0, 0].legend(loc="upper right")

axs[0, 1].plot(
    x_cont,
    cdf_cont,
    linewidth=2,
    label=rf"CDF: $\mathcal{{U}}({a},{b})$",
)

axs[0, 1].set_title("Continuous Uniform CDF")
axs[0, 1].set_xlabel(r"$x$")
axs[0, 1].set_ylabel(r"$\Phi_X(x)=\mathbb{P}(X\leq x)$")
axs[0, 1].set_ylim(-0.05, 1.05)
axs[0, 1].grid(True, linestyle="--", alpha=0.6)
axs[0, 1].legend(loc="lower right")

# Discrete Uniform distribution: fair die
low, high = 1, 6
x_disc = np.arange(1, 7)
pmf_disc = randint.pmf(x_disc, low, high + 1)

axs[1, 0].stem(
    x_disc,
    pmf_disc,
    linefmt="crimson",
    markerfmt="ro",
    basefmt=" ",
)

axs[1, 0].set_title("Discrete Uniform PMF (Fair Die)")
axs[1, 0].set_xlabel(r"$x$")
axs[1, 0].set_ylabel(r"$P_X(x)$")
axs[1, 0].set_xticks(x_disc)
axs[1, 0].set_ylim(-0.02, 0.25)
axs[1, 0].grid(True, linestyle="--", alpha=0.6)

x_disc_cdf = np.arange(0, 8)
cdf_disc = randint.cdf(x_disc_cdf, low, high + 1)

axs[1, 1].step(
    x_disc_cdf,
    cdf_disc,
    where="post",
    color="crimson",
    linewidth=2,
    label="Discrete CDF",
)
axs[1, 1].plot(x_disc_cdf, cdf_disc, "ro", alpha=0.7)

axs[1, 1].set_title("Discrete Uniform CDF")
axs[1, 1].set_xlabel(r"$x$")
axs[1, 1].set_ylabel(r"$\Phi_X(x)=\mathbb{P}(X\leq x)$")
axs[1, 1].set_ylim(-0.05, 1.05)
axs[1, 1].grid(True, linestyle="--", alpha=0.6)
axs[1, 1].legend(loc="lower right")

plt.show()
<Figure size 1200x800 with 4 Axes>

Example: uncertain rotational speed of a mechanical system

Random variables are particularly useful for representing uncertain physical quantities.

Suppose that the rotational speed of a shaft varies during operation. Let

ξ=rotational speed of the shaft (rpm).\xi=\text{rotational speed of the shaft (rpm)}.

Assume that, during a particular operating regime, the rotational speed can take any value between 1800 rpm and 2200 rpm with equal likelihood. We model this by

ξ∼U(1800,2200).\xi\sim U(1800,2200).

Its density is therefore

pξ(x)={1400,1800≤x≤2200,0,otherwise.p_\xi(x) = \begin{cases} \dfrac{1}{400}, & 1800\leq x\leq2200,\\[6pt] 0, & \text{otherwise}. \end{cases}

For example, the probability that the rotational speed lies between 1900 rpm and 2000 rpm is

P({1900≤ξ≤2000})=∫19002000pξ(x) dx=1400∫19002000dx=100400=14.\begin{aligned} \mathbb{P}(\{1900\leq\xi\leq2000\}) &= \int_{1900}^{2000}p_\xi(x)\,dx\\ &= \frac{1}{400}\int_{1900}^{2000}dx\\ &= \frac{100}{400} = \frac14. \end{aligned}

Thus, under the assumed model, there is a 25%25\% probability that the rotational speed lies between 1900 rpm and 2000 rpm.

The same calculation can be performed directly using the CDF:

P({1900≤ξ≤2000})=Φξ(2000)−Φξ(1900)=2000−1800400−1900−1800400=14.\mathbb{P}(\{1900\leq\xi\leq2000\}) = \Phi_\xi(2000)-\Phi_\xi(1900) = \frac{2000-1800}{400} - \frac{1900-1800}{400} = \frac14.

This illustrates the equivalence between the density and CDF descriptions of an absolutely continuous random variable.

Distributions beyond the discrete and continuous cases

The discrete and absolutely continuous cases are the two principal settings considered in this course, but they do not exhaust all possible probability distributions.

For example, a random variable may have both discrete and continuous components. Such a distribution is called a mixed distribution.

There are also continuous distribution functions that cannot be represented by an ordinary probability density function. Thus, the statement that a random variable is continuous should not, in complete generality, be taken to mean that it necessarily has a density.

For the purposes of the present course, however, the distinction between discrete random variables described by PMFs and absolutely continuous random variables described by PDFs will cover the main examples and applications.

Summary

A random variable assigns a real number to each outcome of a random experiment.

For a discrete random variable, the probability distribution is described by the probability mass function

Pξ(x)=P({ξ=x}),P_\xi(x)=\mathbb{P}(\{\xi=x\}),

with

∑xPξ(x)=1.\sum_xP_\xi(x)=1.

For an absolutely continuous random variable, the probability distribution is described by a density pξp_\xi satisfying

pξ(x)≥0,∫−∞∞pξ(x) dx=1.p_\xi(x)\geq0, \qquad \int_{-\infty}^{\infty}p_\xi(x)\,dx=1.

Probabilities are obtained by summation in the discrete case and integration in the continuous case.

The cumulative distribution function

Φξ(x)=P({ξ≤x})\Phi_\xi(x)=\mathbb{P}(\{\xi\leq x\})

provides a unified description of the distribution. In particular,

P({x′<ξ≤x′′})=Φξ(x′′)−Φξ(x′).\mathbb{P}(\{x'<\xi\leq x''\}) = \Phi_\xi(x'')-\Phi_\xi(x').

For a discrete random variable, the CDF is a step function. For a random variable with a density, the CDF is continuous and is obtained by integrating the density.

The next step is to study important families of probability distributions, including the Bernoulli, binomial, Poisson, and normal distributions. These distributions provide models for many common random phenomena and will also provide the foundation for the study of expectation, variance, and limit theorems in the subsequent lectures.