Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Lecture 2 - Probability, Conditional Probability, and Independence

So far we have introduced probability through examples, counting arguments, and relative frequencies. These examples suggest how probabilities should behave, but they do not yet provide a general mathematical definition.

We now introduce the axiomatic definition of probability. This formulation does not require the elementary outcomes to be equally likely and does not depend on a particular interpretation of probability. Once the axioms are established, the usual rules of probability follow from them.

Axiomatic definition of probability

Let Ω\Omega be the sample space of a random experiment, and let F\mathcal{F} denote the collection of events to which we assign probabilities. A probability measure is a function

P:F⟶[0,1]\mathbb{P}:\mathcal{F}\longrightarrow[0,1]

satisfying the following three axioms.

The first two axioms state that probabilities are non-negative and that the probability of an event that is certain to occur is one. The third axiom expresses the idea that probabilities add when events cannot occur simultaneously.

All the elementary rules of probability can be derived from these three properties.

Basic consequences of the axioms

We first determine the probability of the impossible event. Since

Ω=Ω∪∅,\Omega=\Omega\cup\varnothing,

and the two events are mutually exclusive, countable additivity gives

P(Ω)=P(Ω)+P(∅).\mathbb{P}(\Omega) = \mathbb{P}(\Omega)+\mathbb{P}(\varnothing).

Since P(Ω)=1\mathbb{P}(\Omega)=1, it follows that

P(∅)=0.\mathbb{P}(\varnothing)=0.

Next, let AcA^c denote the complementary event of AA. Since AA and AcA^c are mutually exclusive and

A∪Ac=Ω,A\cup A^c=\Omega,

we have

P(A)+P(Ac)=1,\mathbb{P}(A)+\mathbb{P}(A^c)=1,

and therefore

P(Ac)=1−P(A).\mathbb{P}(A^c)=1-\mathbb{P}(A).

The axioms also imply that probability is monotone.

If A⊆BA\subseteq B, then BB can be decomposed into two mutually exclusive events,

B=A∪(B∖A).B=A\cup(B\setminus A).

Hence,

P(B)=P(A)+P(B∖A)≥P(A).\mathbb{P}(B) = \mathbb{P}(A)+\mathbb{P}(B\setminus A) \geq \mathbb{P}(A).

Thus,

A⊆B⟹P(A)≤P(B).A\subseteq B \quad\Longrightarrow\quad \mathbb{P}(A)\leq\mathbb{P}(B).

In particular, since every event is contained in Ω\Omega,

0≤P(A)≤1.0\leq\mathbb{P}(A)\leq1.

Finally, for two arbitrary events AA and BB, we can write

A∪B=A∪(B∖A),A\cup B = A\cup(B\setminus A),

where the two events on the right are mutually exclusive. Since

B=(A∩B)∪(B∖A),B=(A\cap B)\cup(B\setminus A),

we obtain the addition rule

P(A∪B)=P(A)+P(B)−P(A∩B).\mathbb{P}(A\cup B) = \mathbb{P}(A)+\mathbb{P}(B)-\mathbb{P}(A\cap B).

The subtraction of P(A∩B)\mathbb{P}(A\cap B) is necessary because the intersection is counted twice in P(A)+P(B)\mathbb{P}(A)+\mathbb{P}(B).

Conditional probability

The probability of an event describes its likelihood before additional information about the outcome of the experiment is available. In many applications, however, we obtain information about the outcome and want to update the probability accordingly.

Suppose that we know that an event BB has occurred. We are then interested in the probability that another event AA occurs under the condition that BB has occurred.

Provided that P(B)>0\mathbb{P}(B)>0, the conditional probability of AA given BB is defined by

P(A∣B)=P(A∩B)P(B).\mathbb{P}(A\mid B) = \frac{\mathbb{P}(A\cap B)} {\mathbb{P}(B)}.

The event A∩BA\cap B represents the outcomes for which both AA and BB occur. Thus, when we condition on BB, we restrict our attention to the part of the sample space in which BB occurs.

Equiprobable interpretation

The meaning of conditional probability is particularly transparent for a finite experiment with equally likely outcomes. Let NN be the total number of elementary outcomes, N(B)N(B) the number for which BB occurs, and N(A∩B)N(A\cap B) the number for which both AA and BB occur. Then

P(B)=N(B)N,P(A∩B)=N(A∩B)N.\mathbb{P}(B)=\frac{N(B)}{N}, \qquad \mathbb{P}(A\cap B)=\frac{N(A\cap B)}{N}.

Consequently,

P(A∣B)=P(A∩B)P(B)=N(A∩B)N(B).\mathbb{P}(A\mid B) = \frac{\mathbb{P}(A\cap B)} {\mathbb{P}(B)} = \frac{N(A\cap B)} {N(B)}.

Thus, after learning that BB has occurred, we effectively restrict attention to the outcomes satisfying BB. Among these outcomes, the conditional probability is the proportion for which AA also occurs.

This interpretation will be useful when we later discuss statistical independence.

Properties of conditional probability

Conditional probability satisfies the same basic probability rules as ordinary probability.

The definition of conditional probability can also be written as

P(A∩B)=P(A∣B)P(B).\mathbb{P}(A\cap B) = \mathbb{P}(A\mid B)\mathbb{P}(B).

This is often called the multiplication rule. It expresses the probability that both AA and BB occur in terms of the probability of BB and the probability of AA once BB is known.

By interchanging AA and BB, we also have, whenever P(A)>0\mathbb{P}(A)>0,

P(A∩B)=P(B∣A)P(A).\mathbb{P}(A\cap B) = \mathbb{P}(B\mid A)\mathbb{P}(A).

Equating the two expressions gives

P(A∣B)P(B)=P(B∣A)P(A).\mathbb{P}(A\mid B)\mathbb{P}(B) = \mathbb{P}(B\mid A)\mathbb{P}(A).

This relation is the basis of Bayes’ formula.

Bayes’ formula

Suppose that P(A)>0\mathbb{P}(A)>0 and P(B)>0\mathbb{P}(B)>0. From (2) and (3),

P(A∣B)P(B)=P(B∣A)P(A).\mathbb{P}(A\mid B)\mathbb{P}(B) = \mathbb{P}(B\mid A)\mathbb{P}(A).

Therefore,

P(A∣B)=P(B∣A)P(A)P(B).\mathbb{P}(A\mid B) = \frac{ \mathbb{P}(B\mid A)\mathbb{P}(A) }{ \mathbb{P}(B) }.

This is Bayes’ formula. It allows us to reverse the direction of a conditional probability: information about P(B∣A)\mathbb{P}(B\mid A) can be used to determine P(A∣B)\mathbb{P}(A\mid B), provided that the required probabilities are known.

The denominator can often be computed by partitioning the possible causes of BB.

Law of total probability

Suppose that B1,B2,…B_1,B_2,\ldots are mutually exclusive events that form a partition of the sample space:

Bi∩Bj=∅,i≠j,B_i\cap B_j=\varnothing, \qquad i\neq j,

and

⋃k=1∞Bk=Ω.\bigcup_{k=1}^{\infty}B_k=\Omega.

Assume also that P(Bk)>0\mathbb{P}(B_k)>0 for every kk. Then every event AA can be decomposed according to which of the events BkB_k occurs:

A=⋃k=1∞(A∩Bk).A = \bigcup_{k=1}^{\infty}(A\cap B_k).

The events in this union are mutually exclusive. Therefore,

P(A)=∑k=1∞P(A∩Bk).\mathbb{P}(A) = \sum_{k=1}^{\infty} \mathbb{P}(A\cap B_k).

Using the multiplication rule (2), we obtain the law of total probability.

Combining this result with Bayes’ formula gives

P(Bj∣A)=P(A∣Bj)P(Bj)∑k=1∞P(A∣Bk)P(Bk).\mathbb{P}(B_j\mid A) = \frac{ \mathbb{P}(A\mid B_j)\mathbb{P}(B_j) }{ \displaystyle \sum_{k=1}^{\infty} \mathbb{P}(A\mid B_k)\mathbb{P}(B_k) }.

This form is particularly useful when the events BkB_k represent different possible causes of an observed event AA.

Statistical independence

Conditional probability provides a precise way of describing how information about one event changes the probability of another.

Suppose that P(B)>0\mathbb{P}(B)>0. If

P(A∣B)=P(A),\mathbb{P}(A\mid B)=\mathbb{P}(A),

then learning that BB has occurred does not change the probability of AA. This motivates the definition of independence.

The definition is symmetric in AA and BB. Thus, independence is a mutual property: if AA and BB are independent, then BB and AA are also independent.

The connection with conditional probability is immediate. If P(B)>0\mathbb{P}(B)>0, then

P(A∣B)=P(A∩B)P(B).\mathbb{P}(A\mid B) = \frac{\mathbb{P}(A\cap B)} {\mathbb{P}(B)}.

If AA and BB are independent, then

P(A∣B)=P(A)P(B)P(B)=P(A).\mathbb{P}(A\mid B) = \frac{\mathbb{P}(A)\mathbb{P}(B)} {\mathbb{P}(B)} = \mathbb{P}(A).

Conversely, if P(B)>0\mathbb{P}(B)>0 and

P(A∣B)=P(A),\mathbb{P}(A\mid B)=\mathbb{P}(A),

then multiplying by P(B)\mathbb{P}(B) gives

P(A∩B)=P(A)P(B).\mathbb{P}(A\cap B) = \mathbb{P}(A)\mathbb{P}(B).

Therefore,

A and B are independent⟺P(A∣B)=P(A),A\text{ and }B\text{ are independent} \quad\Longleftrightarrow\quad \mathbb{P}(A\mid B)=\mathbb{P}(A),

provided that P(B)>0\mathbb{P}(B)>0.

Similarly, if P(A)>0\mathbb{P}(A)>0,

A and B are independent⟺P(B∣A)=P(B).A\text{ and }B\text{ are independent} \quad\Longleftrightarrow\quad \mathbb{P}(B\mid A)=\mathbb{P}(B).

Thus, independence means precisely that learning that one event has occurred provides no information, in the probabilistic sense, about the other event.

Independence and complements

Independence is preserved when events are replaced by their complements. For example, if AA and BB are independent, then

P(Ac∩B)=P(B)−P(A∩B).\mathbb{P}(A^c\cap B) = \mathbb{P}(B)-\mathbb{P}(A\cap B).

Using (7),

P(Ac∩B)=P(B)−P(A)P(B)=(1−P(A))P(B)=P(Ac)P(B).\begin{aligned} \mathbb{P}(A^c\cap B) &= \mathbb{P}(B) - \mathbb{P}(A)\mathbb{P}(B)\\ &= (1-\mathbb{P}(A))\mathbb{P}(B)\\ &= \mathbb{P}(A^c)\mathbb{P}(B). \end{aligned}

Hence AcA^c and BB are independent. The same argument applies to AA and BcB^c, and therefore also to AcA^c and BcB^c.

Independence of several events

The definition of independence extends naturally to more than two events. For a collection of events, it is not sufficient to require only that every pair be independent.

For a sequence of events A1,A2,…A_1,A_2,\ldots, the same condition is required for every finite collection of distinct events.

This distinction will become important later when we introduce independent random variables. There, independence will mean that the events generated by the different random variables factorize in the corresponding way.

Conditional probability and relative frequency

The definition of conditional probability also gives a useful interpretation in terms of repeated experiments.

Suppose that an experiment is repeated nn times. Let n(A)n(A) denote the number of trials in which AA occurs, and let n(A∩B)n(A\cap B) denote the number in which both AA and BB occur. Among the trials for which BB occurs, the relative frequency of AA is

n(A∩B)n(B),\frac{n(A\cap B)}{n(B)},

provided that n(B)>0n(B)>0.

The definition

P(A∣B)=P(A∩B)P(B)\mathbb{P}(A\mid B) = \frac{\mathbb{P}(A\cap B)} {\mathbb{P}(B)}

therefore corresponds, under the relative-frequency interpretation, to the limiting proportion of occurrences of AA among those trials in which BB occurs.

Similarly, independence means that restricting attention to the trials in which BB occurs does not change the long-run frequency of AA:

P(A∣B)=P(A).\mathbb{P}(A\mid B)=\mathbb{P}(A).

Thus, the axiomatic definition, conditional probability, and the relative-frequency interpretation are consistent with the same underlying probabilistic structure.

Failures in series and parallel systems

Many physical and logical systems can be modeled as assemblies of components arranged in mechanical or logical configurations. Let YiY_i be an indicator variable where Yi=1Y_i = 1 if component ii fails and Yi=0Y_i = 0 if it functions properly, with individual failure probability Pi=P(Yi=1)P_i = \mathbb{P}(Y_i = 1).

Combined System Reliability (Heater, Pumps, and Turbines)

To illustrate how series, parallel, and mm-out-of-nn logic combine when calculating probabilities, consider a power generation facility consisting of three main sub-systems connected in series:

  1. Heater Sub-system: A single heater (R1R_1).

  2. Pump Sub-system: Two pumps (R2R_2 and R3R_3) operating in parallel.

  3. Turbine Sub-system: Five turbines (R4,R5,R6,R7,R8R_4, R_5, R_6, R_7, R_8) operating as a 3-out-of-5 system (requires at least 3 functioning turbines for the sub-system to work).

Combined System Reliability (Heater, Pumps, and Turbines)

Between scheduled maintenances, components fail independently. For simplicity, assume that all turbines have an identical failure probability of 0.15:

ComponentR1R_1 (Heater)R2R_2 (Pump 1)R3R_3 (Pump 2)R4R_4 through R8R_8 (Turbines)
P(Failure)\mathbb{P}(\text{Failure})0.050.100.080.15 each

Let WiW_i denote the event that component ii works properly.


Step 1: Heater Sub-system Reliability

The heater sub-system consists of a single component R1R_1. Its reliability is simply the probability that R1R_1 does not fail:

P(Heater Works)=P(W1)=1−0.05=0.95.\mathbb{P}(\text{Heater Works}) = \mathbb{P}(W_1) = 1 - 0.05 = 0.95.

Step 2: Pump Sub-system Reliability

The pump sub-system operates in parallel, meaning it survives as long as at least one pump works. It fails only if both pumps fail simultaneously:

P(Pumps Work)=1−P(W2c∩W3c).\mathbb{P}(\text{Pumps Work}) = 1 - \mathbb{P}(W_2^c \cap W_3^c).

Since R2R_2 and R3R_3 fail independently:

P(Pumps Work)=1−P(W2c)P(W3c)=1−(0.10)(0.08)=1−0.008=0.992.\mathbb{P}(\text{Pumps Work}) = 1 - \mathbb{P}(W_2^c)\mathbb{P}(W_3^c) = 1 - (0.10)(0.08) = 1 - 0.008 = 0.992.

Step 3: Turbine Sub-system Reliability (mm-out-of-nn System)

The turbine sub-system requires at least 3 out of 5 turbines to function. Because all 5 turbines have identical probabilities of success (1−0.15=0.851 - 0.15 = 0.85) and failure (0.15), any specific configuration with kk working turbines and 5−k5-k failed turbines has probability (0.85)k(0.15)5−k(0.85)^k (0.15)^{5-k}.

Since the different working states are mutually exclusive, we compute the total reliability by counting the number of valid outcome configurations for 5, 4, or 3 working turbines and summing their probabilities:

Summing the probabilities across all 16 mutually exclusive functional states yields:

P(Turbines Work)=0.4437+0.3915+0.1382=0.9734.\mathbb{P}(\text{Turbines Work}) = 0.4437 + 0.3915 + 0.1382 = 0.9734.

Step 4: Overall System Reliability

Because the heater, pump, and turbine sub-systems are connected in series, the overall plant functions if and only if all three sub-systems work. Assuming independence among sub-systems:

P(System Works)=P(Heater Works)×P(Pumps Work)×P(Turbines Work)\mathbb{P}(\text{System Works}) = \mathbb{P}(\text{Heater Works}) \times \mathbb{P}(\text{Pumps Work}) \times \mathbb{P}(\text{Turbines Work})
P(System Works)=(0.95)×(0.992)×(0.9734)≈0.9173.\mathbb{P}(\text{System Works}) = (0.95) \times (0.992) \times (0.9734) \approx 0.9173.

Looking ahead

We have now established the basic rules needed to describe how events interact:

probability⟶conditional probability⟶independence.\text{probability} \longrightarrow \text{conditional probability} \longrightarrow \text{independence}.

In the next lecture, we move from events to random variables. A random variable assigns a numerical value to the outcome of a random experiment, allowing us to describe random phenomena quantitatively. We will introduce probability mass functions, probability densities, and distribution functions, which provide a systematic way of describing the distribution of a random variable.