Lecture 2 - Probability, Conditional Probability, and Independence
So far we have introduced probability through examples, counting arguments, and relative frequencies. These examples suggest how probabilities should behave, but they do not yet provide a general mathematical definition.
We now introduce the axiomatic definition of probability. This formulation does not require the elementary outcomes to be equally likely and does not depend on a particular interpretation of probability. Once the axioms are established, the usual rules of probability follow from them.
Let Ω be the sample space of a random experiment, and let F denote the collection of events to which we assign probabilities. A probability measure is a function
P:F⟶[0,1]
satisfying the following three axioms.
The first two axioms state that probabilities are non-negative and that the probability of an event that is certain to occur is one. The third axiom expresses the idea that probabilities add when events cannot occur simultaneously.
All the elementary rules of probability can be derived from these three properties.
The probability of an event describes its likelihood before additional information about the outcome of the experiment is available. In many applications, however, we obtain information about the outcome and want to update the probability accordingly.
Suppose that we know that an event B has occurred. We are then interested in the probability that another event A occurs under the condition that B has occurred.
Provided that P(B)>0, the conditional probability of A given B is defined by
The event A∩B represents the outcomes for which both A and B occur. Thus, when we condition on B, we restrict our attention to the part of the sample space in which B occurs.
The meaning of conditional probability is particularly transparent for a finite experiment with equally likely outcomes. Let N be the total number of elementary outcomes, N(B) the number for which B occurs, and N(A∩B) the number for which both A and B occur. Then
P(B)=NN(B),P(A∩B)=NN(A∩B).
Consequently,
P(A∣B)=P(B)P(A∩B)=N(B)N(A∩B).
Thus, after learning that B has occurred, we effectively restrict attention to the outcomes satisfying B. Among these outcomes, the conditional probability is the proportion for which A also occurs.
This interpretation will be useful when we later discuss statistical independence.
This is often called the multiplication rule. It expresses the probability that both A and B occur in terms of the probability of B and the probability of A once B is known.
By interchanging A and B, we also have, whenever P(A)>0,
This is Bayes’ formula. It allows us to reverse the direction of a conditional probability: information about P(B∣A) can be used to determine P(A∣B), provided that the required probabilities are known.
The denominator can often be computed by partitioning the possible causes of B.
Thus, independence means precisely that learning that one event has occurred provides no information, in the probabilistic sense, about the other event.
The definition of independence extends naturally to more than two events. For a collection of events, it is not sufficient to require only that every pair be independent.
For a sequence of events A1,A2,…, the same condition is required for every finite collection of distinct events.
This distinction will become important later when we introduce independent random variables. There, independence will mean that the events generated by the different random variables factorize in the corresponding way.
The definition of conditional probability also gives a useful interpretation in terms of repeated experiments.
Suppose that an experiment is repeated n times. Let n(A) denote the number of trials in which A occurs, and let n(A∩B) denote the number in which both A and B occur. Among the trials for which B occurs, the relative frequency of A is
n(B)n(A∩B),
provided that n(B)>0.
The definition
P(A∣B)=P(B)P(A∩B)
therefore corresponds, under the relative-frequency interpretation, to the limiting proportion of occurrences of A among those trials in which B occurs.
Similarly, independence means that restricting attention to the trials in which B occurs does not change the long-run frequency of A:
P(A∣B)=P(A).
Thus, the axiomatic definition, conditional probability, and the relative-frequency interpretation are consistent with the same underlying probabilistic structure.
Many physical and logical systems can be modeled as assemblies of components arranged in mechanical or logical configurations. Let Yi be an indicator variable where Yi=1 if component i fails and Yi=0 if it functions properly, with individual failure probability Pi=P(Yi=1).
Combined System Reliability (Heater, Pumps, and Turbines)¶
To illustrate how series, parallel, and m-out-of-n logic combine when calculating probabilities, consider a power generation facility consisting of three main sub-systems connected in series:
Heater Sub-system: A single heater (R1).
Pump Sub-system: Two pumps (R2 and R3) operating in parallel.
Turbine Sub-system: Five turbines (R4,R5,R6,R7,R8) operating as a 3-out-of-5 system (requires at least 3 functioning turbines for the sub-system to work).
Between scheduled maintenances, components fail independently. For simplicity, assume that all turbines have an identical failure probability of 0.15:
Component
R1 (Heater)
R2 (Pump 1)
R3 (Pump 2)
R4 through R8 (Turbines)
P(Failure)
0.05
0.10
0.08
0.15 each
Let Wi denote the event that component i works properly.
Step 1: Heater Sub-system Reliability
The heater sub-system consists of a single component R1. Its reliability is simply the probability that R1 does not fail:
P(Heater Works)=P(W1)=1−0.05=0.95.
Step 2: Pump Sub-system Reliability
The pump sub-system operates in parallel, meaning it survives as long as at least one pump works. It fails only if both pumps fail simultaneously:
The turbine sub-system requires at least 3 out of 5 turbines to function. Because all 5 turbines have identical probabilities of success (1−0.15=0.85) and failure (0.15), any specific configuration with k working turbines and 5−k failed turbines has probability (0.85)k(0.15)5−k.
Since the different working states are mutually exclusive, we compute the total reliability by counting the number of valid outcome configurations for 5, 4, or 3 working turbines and summing their probabilities:
Case 1: All 5 turbines work ((55)=1 state)
1×(0.85)5=0.4437.
Case 2: Exactly 4 turbines work ((45)=5 states)
5×(0.85)4(0.15)1=5×(0.5220)(0.15)=0.3915.
Case 3: Exactly 3 turbines work ((35)=10 states)
10×(0.85)3(0.15)2=10×(0.6141)(0.0225)=0.1382.
Summing the probabilities across all 16 mutually exclusive functional states yields:
P(Turbines Work)=0.4437+0.3915+0.1382=0.9734.
Step 4: Overall System Reliability
Because the heater, pump, and turbine sub-systems are connected in series, the overall plant functions if and only if all three sub-systems work. Assuming independence among sub-systems:
We have now established the basic rules needed to describe how events interact:
probability⟶conditional probability⟶independence.
In the next lecture, we move from events to random variables. A random variable assigns a numerical value to the outcome of a random experiment, allowing us to describe random phenomena quantitatively. We will introduce probability mass functions, probability densities, and distribution functions, which provide a systematic way of describing the distribution of a random variable.