Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

The best way to begin understanding probability is through a concrete example. Consider the simplest experiment of all: tossing an unbiased coin. This experiment has only two mutually exclusive outcomes:

  1. Heads (HH).

  2. Tails (TT).

Because the factors influencing the outcome are far too complex and numerous to model deterministically, we regard the result as random.

What is the probability of each outcome? Since the coin is unbiased, both outcomes are equally likely. We therefore assign probability 12\frac{1}{2} to each outcome.

Generalizing this idea to any experiment with a finite number of mutually exclusive and equally likely outcomes, we say that the outcomes are equiprobable and obtain the following classical definition.


The Frequentist Approach

The frequentist approach interprets probability in terms of the empirical behavior of relative frequencies observed over repeated trials. Under this viewpoint, we imagine that an experiment can be repeated indefinitely under identical conditions.

Suppose that an experiment is repeated nn times in a sequence of independent trials. If n(A)n(A) denotes the number of trials in which event AA occurs, then the ratio

f(A)=n(A)nf(A)=\frac{n(A)}{n}

is called the relative frequency of the event AA.

The frequentist interpretation identifies the probability of AA with the limiting relative frequency, whenever this limit exists:

P(A)=lim⁡n→∞n(A)n.\mathbb{P}(A) = \lim_{n\to\infty}\frac{n(A)}{n}.

This interpretation leads naturally to the need to count possible and favorable outcomes. Such questions belong to the field of combinatorics.

Counting Things

Combinatorics encompasses a large collection of techniques for counting finite sets. We will use a few fundamental principles that are particularly useful in probability.


Fundamental Principles of Counting

This leads to the following fundamental rule.

Sampling Strategies

When drawing rr objects from a population of nn distinct objects, the number of possible arrangements depends on whether replacement is allowed and whether order matters.

Combinations and Partitions

We introduce the following notation for this quantity.

The Sample Space

We call the mutually exclusive outcomes of a random experiment elementary events (or sample points), and denote them by the Greek letter ω\omega.

The set of all possible elementary events associated with a given experiment is called the sample space, denoted by Ω\Omega.

An event AA associated with Ω\Omega is a set of elementary outcomes such that, for every ω∈Ω\omega\in\Omega, we can unambiguously determine whether ω\omega leads to the occurrence of AA. We therefore identify an event with the corresponding subset of the sample space.

Special cases of events include:

Given two events A1A_1 and A2A_2:

A1∩A2=∅.A_1\cap A_2=\emptyset.

We use the standard set operations on events:

⋃kAk.\bigcup_k A_k.
⋂kAk.\bigcap_k A_k.
A‾=Ω∖A.\overline{A}=\Omega\setminus A.

Constructing a Probability Model

A random experiment by itself does not determine a probability distribution. To describe uncertainty mathematically, we need to specify not only what outcomes are possible, but also how probability is assigned to those outcomes.

This leads to the distinction between three fundamental ingredients:

random experiment⟶sample space⟶probability measure.\boxed{ \text{random experiment} \quad\longrightarrow\quad \text{sample space} \quad\longrightarrow\quad \text{probability measure}. }

The random experiment is the physical, computational, or conceptual procedure whose outcome is uncertain. The sample space Ω\Omega specifies the set of all possible outcomes of the experiment. The probability measure P\mathbb P assigns probabilities to the events that can occur.

Together, these objects form a probability model,

(Ω,F,P),(\Omega,\mathcal F,\mathbb P),

where F\mathcal F is the collection of events to which probabilities are assigned.

The distinction is important because specifying the sample space alone does not specify how likely its outcomes are.

Probability models and physical experiments

In applications, the probability measure is usually not obtained from the experiment alone. It is a mathematical model based on available information, experimental data, physical considerations, or simplifying assumptions.

For example, consider the lifetime of a mechanical component. The experiment consists of selecting a component and observing how long it operates before failure. A natural sample space is

Ω=(0,∞),\Omega=(0,\infty),

since the lifetime is a positive quantity.

An event might be

A={ω∈Ω:ω>1000},A=\{\omega\in\Omega:\omega>1000\},

corresponding to the component surviving more than 1000 hours.

The sample space tells us which lifetimes are possible, but it does not determine the probability

P(A).\mathbb P(A).

To obtain this probability, we need a model for the distribution of component lifetimes. Such a model could be based on reliability experiments or on a physical theory of the failure mechanism.

This distinction is particularly important in engineering and computational applications. The mathematical model is an idealization of a physical system, and its usefulness depends on how well it represents the features of the experiment that are relevant to the question being studied.

Different models for the same experiment

The same physical experiment can sometimes be represented by different probability models, depending on the level of detail that is relevant.

For example, suppose that the experiment consists of measuring the diameter of a manufactured shaft.

At one level, we might record the actual diameter,

Ω=(0,∞).\Omega=(0,\infty).

At another level, if the manufacturing process only distinguishes whether the shaft satisfies a tolerance specification, we could use the simpler sample space

Ω={within tolerance,out of tolerance}.\Omega=\{\text{within tolerance},\text{out of tolerance}\}.

Both descriptions refer to the same physical process, but they answer different questions.

The first model retains quantitative information about the diameter, while the second records only whether the component satisfies the required specification.

Thus, constructing a probability model involves deciding which aspects of the random experiment need to be represented.

Discrete and continuous models

The nature of the sample space also depends on the experiment.

For a finite experiment, the sample space may contain only finitely many outcomes, such as

Ω={1,2,…,6}\Omega=\{1,2,\ldots,6\}

for a die.

A countably infinite sample space can arise when the possible outcomes are

Ω={0,1,2,…}.\Omega=\{0,1,2,\ldots\}.

For example, an outcome might represent the number of trials required before a particular event occurs.

In other experiments, the possible outcomes form an interval or another uncountable set. For example, a measured time or length might take values in

Ω=[0,1].\Omega=[0,1].

These different types of sample spaces require different ways of specifying the probability measure. In particular, when the sample space is continuous, individual outcomes can have probability zero even though the corresponding outcomes are possible.

The distinction between discrete and continuous probability models will become central when we introduce random variables and their probability distributions.

From an experiment to a mathematical model

The construction of a probability model can therefore be viewed as a sequence of questions:

What is the random experiment?What are the possible outcomes?Which collections of outcomes are events?How should probability be assigned to those events?\boxed{ \begin{aligned} &\text{What is the random experiment?}\\ &\text{What are the possible outcomes?}\\ &\text{Which collections of outcomes are events?}\\ &\text{How should probability be assigned to those events?} \end{aligned} }

The first three questions determine the structure of the model, while the last specifies the probability measure.

A well-chosen probability model should capture the aspects of the experiment that are relevant to the problem under consideration. Once the model has been specified, the axioms of probability provide the rules that all probability calculations must satisfy.

In the following lectures, we will develop additional tools for working within such models. In particular, we will see how probabilities change when additional information is available, leading naturally to the concepts of conditional probability, Bayes’ theorem, and independence.

From the Algebra of Events to the Computation of Probabilities

We now derive the basic algebraic rules satisfied by probabilities.

We assume that all events under consideration have well-defined probabilities. Moreover, all events obtained from a given collection of events by taking unions, intersections, differences, and complements are also assumed to have well-defined probabilities.

Consider two mutually exclusive events A1A_1 and A2A_2, and let

A=A1∪A2.A=A_1\cup A_2.

Suppose that we repeat the experiment a large number of times under identical conditions. Let nn be the total number of trials, and let n(A1)n(A_1), n(A2)n(A_2), and n(A)n(A) denote the numbers of trials in which A1A_1, A2A_2, and AA occur, respectively.

Because A1A_1 and A2A_2 are mutually exclusive, whenever AA occurs, exactly one of A1A_1 or A2A_2 occurs. Therefore,

n(A)=n(A1)+n(A2),n(A)=n(A_1)+n(A_2),

and hence

n(A)n=n(A1)n+n(A2)n.\frac{n(A)}{n} = \frac{n(A_1)}{n} + \frac{n(A_2)}{n}.

Passing to the corresponding probabilities gives

P(A)=P(A1)+P(A2).\mathbb{P}(A) = \mathbb{P}(A_1)+\mathbb{P}(A_2).

Similarly, if A1A_1, A2A_2, and A3A_3 are mutually exclusive, then

P(A1∪A2∪A3)=P(A1)+P(A2)+P(A3).\mathbb{P}(A_1\cup A_2\cup A_3) = \mathbb{P}(A_1) + \mathbb{P}(A_2) + \mathbb{P}(A_3).

More generally, for nn mutually exclusive events A1,…,AnA_1,\ldots,A_n, we obtain the addition law for probabilities:

P(⋃k=1nAk)=∑k=1nP(Ak).\mathbb{P}\left(\bigcup_{k=1}^n A_k\right) = \sum_{k=1}^n\mathbb{P}(A_k).

More generally, Theorem 5 can be used to consider the case in which A1,…,AnA_1,\ldots,A_n are pairwise disjoint events, then

P(⋃i=1nAi)=∑i=1nP(Ai).\mathbb P\left(\bigcup_{i=1}^n A_i\right) = \sum_{i=1}^n\mathbb P(A_i).

Finite additivity is particularly useful when an event can be decomposed into simpler, mutually exclusive events.

If two events are not disjoint, we cannot simply add their probabilities. Indeed, if AA and BB overlap, then the outcomes in A∩BA\cap B would be counted twice:

P(A)+P(B).\mathbb P(A)+\mathbb P(B).

To correct this double counting, we subtract the probability of the intersection using the complete formula from Theorem 5.

Inclusion-Exclusion Principle

The addition law can be generalized to events that are not mutually exclusive.

Infinite Sequences and Continuity Properties

The addition law for disjoint events is not restricted to a finite number of events. In many probability models, however, it is natural to consider a sequence of events

A1,A2,A3,…A_1,A_2,A_3,\ldots

rather than only a finite collection.

This occurs, for example, when an experiment can produce infinitely many possible outcomes, or when we are interested in whether an event eventually occurs during a sequence of trials. Consider an experiment that is repeated indefinitely and let AkA_k denote the event that a specified outcome occurs for the first time on the kk-th trial. The event that the outcome occurs at some point during the experiment is then

⋃k=1∞Ak.\bigcup_{k=1}^{\infty}A_k.

If the events A1,A2,…A_1,A_2,\ldots are mutually exclusive, exactly one of them can occur. We would therefore expect the probability of their union to be obtained by adding their individual probabilities, just as in the finite case.

This leads naturally to the question of whether the addition law should continue to hold when the number of events is infinite. The answer is one of the fundamental principles of probability: probability is countably additive.

More precisely, if A1,A2,…A_1,A_2,\ldots are pairwise disjoint events, then

P(⋃k=1∞Ak)=∑k=1∞P(Ak).\mathbb{P}\left(\bigcup_{k=1}^{\infty}A_k\right) = \sum_{k=1}^{\infty}\mathbb{P}(A_k).

This is the countable additivity property of probability.

It is important to distinguish this statement from the finite addition law. For any finite collection of pairwise disjoint events,

P(⋃k=1nAk)=∑k=1nP(Ak),\mathbb{P}\left(\bigcup_{k=1}^{n}A_k\right) = \sum_{k=1}^{n}\mathbb{P}(A_k),

whereas countable additivity concerns the limit of such finite unions:

⋃k=1∞Ak=lim⁡n→∞⋃k=1nAk.\bigcup_{k=1}^{\infty}A_k = \lim_{n\to\infty} \bigcup_{k=1}^{n}A_k.

Thus, countable additivity allows the probability model to remain consistent when we pass from finite collections of events to infinite sequences of mutually exclusive events.

This property is essential because infinite sequences arise naturally in probability. Repeated experiments, waiting times, random walks, convergence questions, and many other probabilistic constructions require us to consider what happens after arbitrarily many trials. Countable additivity provides the mathematical rule that connects these infinite constructions with the probabilities of the individual events.

Countable additivity gives us a way to compute the probability of a countable union of disjoint events. In applications, however, we often encounter sequences of events that are not disjoint but instead become progressively smaller or more restrictive.

For example, suppose that AnA_n represents the event that an experiment satisfies a certain condition after nn successive tests. As more tests are performed, the condition may become increasingly restrictive, leading to a decreasing sequence

A1⊇A2⊇A3⊇⋯ .A_1\supseteq A_2\supseteq A_3\supseteq\cdots.

The event that satisfies the condition at every stage, no matter how far the sequence is continued, is

⋂k=1∞Ak.\bigcap_{k=1}^{\infty}A_k.

It is therefore natural to ask whether the probability of this limiting event can be obtained by taking the limit of the probabilities of the successive events:

P(⋂k=1∞Ak)=?lim⁡n→∞P(An).\mathbb{P}\left(\bigcap_{k=1}^{\infty}A_k\right) \stackrel{?}{=} \lim_{n\to\infty}\mathbb{P}(A_n).

This question is the event-theoretic analogue of a familiar idea from analysis: when a sequence of sets becomes smaller and smaller and converges to a limiting set, we would like the corresponding probabilities to converge to the probability of that limiting set.

For decreasing sequences of events, this is indeed the case.

The result is called continuity from above because the events decrease as the index increases, while their probabilities converge to the probability of their limiting intersection.

This property will be useful whenever an event is described as satisfying an infinite sequence of increasingly restrictive conditions. It also illustrates an important feature of a probability measure: countable additivity is not only a rule for adding probabilities, but also provides a consistent notion of taking limits of events.


The inclusion–exclusion principle gives an exact expression for the probability of a finite union of events by correcting for overlaps. For two events,

P(A∪B)=P(A)+P(B)−P(A∩B).\mathbb{P}(A\cup B) = \mathbb{P}(A)+\mathbb{P}(B)-\mathbb{P}(A\cap B).

The correction term is necessary because adding P(A)\mathbb{P}(A) and P(B)\mathbb{P}(B) counts the outcomes in A∩BA\cap B twice.

For a larger collection of events, inclusion–exclusion requires increasingly complicated intersections. In particular, for a countably infinite sequence

A1,A2,…,A_1,A_2,\ldots,

an exact inclusion–exclusion formula would involve intersections of two events, three events, four events, and so on.

In many applications, however, we do not need an exact value. It is enough to have a simple upper bound for the probability that at least one of the events occurs.

The basic idea is straightforward: if we add the probabilities

P(A1)+P(A2)+⋯ ,\mathbb{P}(A_1)+\mathbb{P}(A_2)+\cdots,

any outcome belonging to several of the events may be counted more than once. Therefore, the sum can overestimate the probability of the union, but it cannot underestimate it.

This leads to the following fundamental bound.

The result is also called the union bound. It is particularly useful because it does not require the events to be disjoint or independent, and it avoids the need to compute any intersection probabilities.

For a finite collection, the same idea gives

P(⋃k=1nAk)⩽∑k=1nP(Ak).\mathbb{P}\left(\bigcup_{k=1}^{n}A_k\right) \leqslant \sum_{k=1}^{n}\mathbb{P}(A_k).

Thus, the probability of at least one event occurring is never larger than the sum of the individual probabilities.

The inequality becomes an equality when the events are pairwise disjoint, since in that case there is no multiple counting. When the events overlap, the sum generally overestimates the probability of the union.

The importance of Boole’s inequality goes beyond its simple proof. It provides a way of controlling complicated events by decomposing them into simpler events. This idea will recur throughout probability theory: rather than computing a probability exactly, we can often obtain a useful bound by covering the event of interest with a collection of simpler events.