The best way to begin understanding probability is through a concrete example. Consider the simplest experiment of all: tossing an unbiased coin. This experiment has only two mutually exclusive outcomes:
Heads (H).
Tails (T).
Because the factors influencing the outcome are far too complex and numerous to model deterministically, we regard the result as random.
What is the probability of each outcome? Since the coin is unbiased, both outcomes are equally likely. We therefore assign probability 21 to each outcome.
Generalizing this idea to any experiment with a finite number of mutually exclusive and equally likely outcomes, we say that the outcomes are equiprobable and obtain the following classical definition.
The frequentist approach interprets probability in terms of the empirical behavior of relative frequencies observed over repeated trials. Under this viewpoint, we imagine that an experiment can be repeated indefinitely under identical conditions.
Suppose that an experiment is repeated n times in a sequence of independent trials. If n(A) denotes the number of trials in which event A occurs, then the ratio
f(A)=nn(A)
is called the relative frequency of the event A.
The frequentist interpretation identifies the probability of A with the limiting relative frequency, whenever this limit exists:
Combinatorics encompasses a large collection of techniques for counting finite sets. We will use a few fundamental principles that are particularly useful in probability.
When drawing r objects from a population of n distinct objects, the number of possible arrangements depends on whether replacement is allowed and whether order matters.
Sampling with replacement (ordered).
Each draw has n possibilities, so the number of distinct ordered samples is
N=nr.
Sampling without replacement (ordered).
Each successive draw reduces the number of available choices by one. Therefore, the number of ordered samples, also called r-permutations, is
N=n(n−1)(n−2)⋯(n−r+1)=(n−r)!n!.
Special case. If r=n, this gives the total number of permutations of n objects:
N=n!.
Combinations and Partitions
We introduce the following notation for this quantity.
We call the mutually exclusive outcomes of a random experiment elementary events (or sample points), and denote them by the Greek letter ω.
The set of all possible elementary events associated with a given experiment is called the sample space, denoted by Ω.
An event A associated with Ω is a set of elementary outcomes such that, for every ω∈Ω, we can unambiguously determine whether ω leads to the occurrence of A. We therefore identify an event with the corresponding subset of the sample space.
Special cases of events include:
The sure (or certain) event: this event always occurs, regardless of the outcome of the experiment. It is the entire sample space Ω.
The impossible event: this event never occurs and is represented by the empty set ∅.
Operations on Events
Given two events A1 and A2:
Equivalence:A1 and A2 are identical, or equivalent, if they occur under exactly the same circumstances. We write A1=A2.
Mutual exclusivity:A1 and A2 are mutually exclusive (or incompatible) if they cannot occur simultaneously, that is,
A1∩A2=∅.
We use the standard set operations on events:
Union:A1∪A2 is the event that at least one of A1 or A2 occurs. For a collection A1,A2,…, we write
k⋃Ak.
Intersection:A1∩A2 is the event that both A1 and A2 occur. For a collection A1,A2,…, we write
k⋂Ak.
Difference:A1∖A2 is the event that A1 occurs but A2 does not.
Complement: the complement of A, denoted by A or Ac, is the event that A does not occur. It is given by
A random experiment by itself does not determine a probability distribution. To describe uncertainty mathematically, we need to specify not only what outcomes are possible, but also how probability is assigned to those outcomes.
This leads to the distinction between three fundamental ingredients:
random experiment⟶sample space⟶probability measure.
The random experiment is the physical, computational, or conceptual procedure whose outcome is uncertain. The sample spaceΩ specifies the set of all possible outcomes of the experiment. The probability measureP assigns probabilities to the events that can occur.
Together, these objects form a probability model,
(Ω,F,P),
where F is the collection of events to which probabilities are assigned.
The distinction is important because specifying the sample space alone does not specify how likely its outcomes are.
In applications, the probability measure is usually not obtained from the experiment alone. It is a mathematical model based on available information, experimental data, physical considerations, or simplifying assumptions.
For example, consider the lifetime of a mechanical component. The experiment consists of selecting a component and observing how long it operates before failure. A natural sample space is
Ω=(0,∞),
since the lifetime is a positive quantity.
An event might be
A={ω∈Ω:ω>1000},
corresponding to the component surviving more than 1000 hours.
The sample space tells us which lifetimes are possible, but it does not determine the probability
P(A).
To obtain this probability, we need a model for the distribution of component lifetimes. Such a model could be based on reliability experiments or on a physical theory of the failure mechanism.
This distinction is particularly important in engineering and computational applications. The mathematical model is an idealization of a physical system, and its usefulness depends on how well it represents the features of the experiment that are relevant to the question being studied.
The same physical experiment can sometimes be represented by different probability models, depending on the level of detail that is relevant.
For example, suppose that the experiment consists of measuring the diameter of a manufactured shaft.
At one level, we might record the actual diameter,
Ω=(0,∞).
At another level, if the manufacturing process only distinguishes whether the shaft satisfies a tolerance specification, we could use the simpler sample space
Ω={within tolerance,out of tolerance}.
Both descriptions refer to the same physical process, but they answer different questions.
The first model retains quantitative information about the diameter, while the second records only whether the component satisfies the required specification.
Thus, constructing a probability model involves deciding which aspects of the random experiment need to be represented.
The nature of the sample space also depends on the experiment.
For a finite experiment, the sample space may contain only finitely many outcomes, such as
Ω={1,2,…,6}
for a die.
A countably infinite sample space can arise when the possible outcomes are
Ω={0,1,2,…}.
For example, an outcome might represent the number of trials required before a particular event occurs.
In other experiments, the possible outcomes form an interval or another uncountable set. For example, a measured time or length might take values in
Ω=[0,1].
These different types of sample spaces require different ways of specifying the probability measure. In particular, when the sample space is continuous, individual outcomes can have probability zero even though the corresponding outcomes are possible.
The distinction between discrete and continuous probability models will become central when we introduce random variables and their probability distributions.
The construction of a probability model can therefore be viewed as a sequence of questions:
What is the random experiment?What are the possible outcomes?Which collections of outcomes are events?How should probability be assigned to those events?
The first three questions determine the structure of the model, while the last specifies the probability measure.
A well-chosen probability model should capture the aspects of the experiment that are relevant to the problem under consideration. Once the model has been specified, the axioms of probability provide the rules that all probability calculations must satisfy.
In the following lectures, we will develop additional tools for working within such models. In particular, we will see how probabilities change when additional information is available, leading naturally to the concepts of conditional probability, Bayes’ theorem, and independence.
From the Algebra of Events to the Computation of Probabilities¶
We now derive the basic algebraic rules satisfied by probabilities.
We assume that all events under consideration have well-defined probabilities. Moreover, all events obtained from a given collection of events by taking unions, intersections, differences, and complements are also assumed to have well-defined probabilities.
Consider two mutually exclusive events A1 and A2, and let
A=A1∪A2.
Suppose that we repeat the experiment a large number of times under identical conditions. Let n be the total number of trials, and let n(A1), n(A2), and n(A) denote the numbers of trials in which A1, A2, and A occur, respectively.
Because A1 and A2 are mutually exclusive, whenever A occurs, exactly one of A1 or A2 occurs. Therefore,
n(A)=n(A1)+n(A2),
and hence
nn(A)=nn(A1)+nn(A2).
Passing to the corresponding probabilities gives
P(A)=P(A1)+P(A2).
Similarly, if A1, A2, and A3 are mutually exclusive, then
P(A1∪A2∪A3)=P(A1)+P(A2)+P(A3).
More generally, for n mutually exclusive events A1,…,An, we obtain the addition law for probabilities:
P(k=1⋃nAk)=k=1∑nP(Ak).
More generally, Theorem 5 can be used to consider the case in which
A1,…,An are pairwise disjoint events, then
The addition law for disjoint events is not restricted to a finite number of events. In many probability models, however, it is natural to consider a sequence of events
A1,A2,A3,…
rather than only a finite collection.
This occurs, for example, when an experiment can produce infinitely many possible outcomes, or when we are interested in whether an event eventually occurs during a sequence of trials. Consider an experiment that is repeated indefinitely and let Ak denote the event that a specified outcome occurs for the first time on the k-th trial. The event that the outcome occurs at some point during the experiment is then
k=1⋃∞Ak.
If the events A1,A2,… are mutually exclusive, exactly one of them can occur. We would therefore expect the probability of their union to be obtained by adding their individual probabilities, just as in the finite case.
This leads naturally to the question of whether the addition law should continue to hold when the number of events is infinite. The answer is one of the fundamental principles of probability: probability is countably additive.
More precisely, if A1,A2,… are pairwise disjoint events, then
P(k=1⋃∞Ak)=k=1∑∞P(Ak).
This is the countable additivity property of probability.
It is important to distinguish this statement from the finite addition law. For any finite collection of pairwise disjoint events,
P(k=1⋃nAk)=k=1∑nP(Ak),
whereas countable additivity concerns the limit of such finite unions:
k=1⋃∞Ak=n→∞limk=1⋃nAk.
Thus, countable additivity allows the probability model to remain consistent when we pass from finite collections of events to infinite sequences of mutually exclusive events.
This property is essential because infinite sequences arise naturally in probability. Repeated experiments, waiting times, random walks, convergence questions, and many other probabilistic constructions require us to consider what happens after arbitrarily many trials. Countable additivity provides the mathematical rule that connects these infinite constructions with the probabilities of the individual events.
Countable additivity gives us a way to compute the probability of a countable union of disjoint events. In applications, however, we often encounter sequences of events that are not disjoint but instead become progressively smaller or more restrictive.
For example, suppose that An represents the event that an experiment satisfies a certain condition after n successive tests. As more tests are performed, the condition may become increasingly restrictive, leading to a decreasing sequence
A1⊇A2⊇A3⊇⋯.
The event that satisfies the condition at every stage, no matter how far the sequence is continued, is
k=1⋂∞Ak.
It is therefore natural to ask whether the probability of this limiting event can be obtained by taking the limit of the probabilities of the successive events:
P(k=1⋂∞Ak)=?n→∞limP(An).
This question is the event-theoretic analogue of a familiar idea from analysis: when a sequence of sets becomes smaller and smaller and converges to a limiting set, we would like the corresponding probabilities to converge to the probability of that limiting set.
For decreasing sequences of events, this is indeed the case.
The result is called continuity from above because the events decrease as the index increases, while their probabilities converge to the probability of their limiting intersection.
This property will be useful whenever an event is described as satisfying an infinite sequence of increasingly restrictive conditions. It also illustrates an important feature of a probability measure: countable additivity is not only a rule for adding probabilities, but also provides a consistent notion of taking limits of events.
The inclusion–exclusion principle gives an exact expression for the probability of a finite union of events by correcting for overlaps. For two events,
P(A∪B)=P(A)+P(B)−P(A∩B).
The correction term is necessary because adding P(A) and P(B) counts the outcomes in A∩B twice.
For a larger collection of events, inclusion–exclusion requires increasingly complicated intersections. In particular, for a countably infinite sequence
A1,A2,…,
an exact inclusion–exclusion formula would involve intersections of two events, three events, four events, and so on.
In many applications, however, we do not need an exact value. It is enough to have a simple upper bound for the probability that at least one of the events occurs.
The basic idea is straightforward: if we add the probabilities
P(A1)+P(A2)+⋯,
any outcome belonging to several of the events may be counted more than once. Therefore, the sum can overestimate the probability of the union, but it cannot underestimate it.
This leads to the following fundamental bound.
The result is also called the union bound. It is particularly useful because it does not require the events to be disjoint or independent, and it avoids the need to compute any intersection probabilities.
For a finite collection, the same idea gives
P(k=1⋃nAk)⩽k=1∑nP(Ak).
Thus, the probability of at least one event occurring is never larger than the sum of the individual probabilities.
The inequality becomes an equality when the events are pairwise disjoint, since in that case there is no multiple counting. When the events overlap, the sum generally overestimates the probability of the union.
The importance of Boole’s inequality goes beyond its simple proof. It provides a way of controlling complicated events by decomposing them into simpler events. This idea will recur throughout probability theory: rather than computing a probability exactly, we can often obtain a useful bound by covering the event of interest with a collection of simpler events.