In this chapter we introduce conditions that affect the occurrence of events. Conditional probability is a way to update the probability of an event once we know that another event has occurred. The same idea appears in statistics, finance, medicine, and the sciences, wherever likelihoods have to be evaluated against known information.

The chapter will begin with a rigorous definition of conditional probability, followed by the derivation of the formula used to calculate it. We will explore examples that illustrate how conditional probability applies in various real-world scenarios, helping to clarify this seemingly abstract concept.

Additionally, we will discuss the concept of independence of events. Independence is a key notion that simplifies the computation of probabilities in situations where multiple events occur simultaneously. Understanding when events are independent, and when they are not, is crucial for correctly applying probabilistic models to real-life problems.

We will cover the following key topics in this chapter:

  • The definition and intuition behind conditional probability.

  • The formula for conditional probability and how it is derived from the fundamental principles of probability.

  • The Multiplication Rule for calculating the probability of the intersection of two events, based on conditional probability.

  • The concept of independence, how it can be recognized, and its implications for the calculation of probabilities.

  • Bayes's Theorem, a pivotal result in probability that relies on the concept of conditional probability to reverse conditional relationships.

  • Practical applications and examples that demonstrate the use and impact of conditional probability and independence in various fields.

 Through the exploration of these topics, the chapter aims to provide a comprehensive understanding of how probabilities are adjusted based on conditions and dependencies between events, equipping you with the analytical tools necessary for tackling complex probabilistic scenarios. We will conclude the chapter with exercises that reinforce these concepts and challenge you to apply what you have learned in theoretical and practical contexts.

Conditional Probability

We first introduce probability with conditions. Like its verbatim meaning, conditional probabilities are obtained by adding constraints to existing probability. Essentially, when we are trying to find conditional probability, we are looking for a new probability in from a compressed sample space.

Basic Conditional Probability

DefinitionConditional probability

Given two events AA and BB in a probability space with P(B)>0P(B) > 0, the conditional probability of AA given BB is

P(A∣B)=P(A∩B)P(B).P(A \mid B) = \frac{P(A \cap B)}{P(B)}.
Remark

This definition assumes that P(B)>0P(B) > 0, because conditional probability is undefined when P(B)=0P(B) = 0. The definition essentially describes how the likelihood of AA is adjusted by the occurrence of BB, providing a way to update beliefs about the likelihood of an event based on new information about another event.

Proof

We need to demonstrate that the definition of P(E∣F)P(E \mid F) as P(EF)P(F)\frac{P(EF)}{P(F)} is logically consistent and adheres to the axioms of probability.

Let's consider the sample space Ω\Omega, where F⊆ΩF \subseteq \Omega and P(F)>0P(F) > 0. By definition, EFEF is the event where both EE and FF occur. We want to compute P(E∣F)P(E \mid F), the probability of EE occurring given that FF has occurred.

Step 1: Relative Frequency Interpretation In the relative frequency interpretation of probability, if we were to repeat an experiment a large number of times, P(F)P(F) would be the proportion of times that event FF occurs, and P(EF)P(EF) would be the proportion of times both events EE and FF occur together.

Step 2: Conditional Probability Calculation Given that FF has occurred, the new sample space is restricted to FF. Hence, we need to adjust all probabilities to this new sample space. The probability of any event EE in this new sample space (i.e., under the condition FF) is the proportion of FF that also belongs to EE, which is exactly EFEF. Thus,

P(E∣F)=Number of outcomes in EFNumber of outcomes in F=P(EF)P(F)P(E \mid F) = \frac{\text{Number of outcomes in } EF}{\text{Number of outcomes in } F} = \frac{P(EF)}{P(F)}

Therefore, the definition P(E∣F)=P(EF)P(F)P(E \mid F) = \frac{P(EF)}{P(F)} is a valid and logical extension of the probability measure to the conditioned space where FF has occurred.

We illustrate this with an example from dice again.

ExampleRolling a die twice

Consider the experiment of rolling a fair six-sided die twice. We are interested in the conditional probability of the sum of the numbers on the two dice being greater than 8, given that the first die shows a 4.

The total sample space can be drawn as a lattice, where each point (x,y)(x,y) is the outcome with first roll xx and second roll yy. Red marks are the outcomes whose sum is greater than 8.

Lattice of two dice rolls, with sums greater than 8 marked.

Let AA be the event that the sum is greater than 8, and let BB be the event that the first roll is 4. Conditioning on BB restricts attention to the fourth column. The outcomes in A∩BA \cap B are (4,5)(4,5) and (4,6)(4,6).

The column of outcomes with first roll equal to 4.

Then P(A∩B)=236=118P(A \cap B) = \frac{2}{36} = \frac{1}{18} and P(B)=636=16P(B) = \frac{6}{36} = \frac{1}{6}, so

P(A∣B)=P(A∩B)P(B)=1/181/6=13.P(A \mid B) = \frac{P(A \cap B)}{P(B)} = \frac{1/18}{1/6} = \frac{1}{3}.

This is the same as 26=13\frac{2}{6}=\frac{1}{3} counted inside the subspace BB.

Remark

Both readings are valid: one uses the original measure, the other counts directly in the compressed sample space. The second is easier when the subspace cannot be drawn.

With this example, I believe that you must know how conditional probability is all about. However, this is only a very special case. You may have noticed that for rolling a fair dice, we have equal possibility to get 1-6 in each roll. In the real life, most events have different probabilities. Let's see another example of unfaired dice. But to solve this problem, we need to use an important conclusion of conditional probability.

Corollary

By multiplying P(F)P(F) to P(E∣F)=P(EF)P(F)P(E \mid F) = \frac{P(EF)}{P(F)}, we find that P(EF)=P(E∣F)P(F)P(EF) = P(E \mid F)P(F)

Example

An unfair four-sided die (numbered 1 to 4) is rolled twice. The biases for the die are:

  • P(1)=0.1P(1) = 0.1

  • P(2)=0.2P(2) = 0.2

  • P(3)=0.3P(3) = 0.3

  • P(4)=0.4P(4) = 0.4

We are tasked with calculating:

  1. The conditional probability that the sum of the two rolls is 4, given that the first roll is a 2.

  2. The conditional probability that the first roll is a 2, given that the sum of the two rolls is 4.

Solution

Define the events:

  • AA: The sum of the two rolls is 4.

  • BB: The first roll is a 2.

We can draw a tree diagram to show all possible results with the probability of each branch.

Tree of two rolls of an unfair four-sided die.

1. P(A∣B)P(A \mid B)

The event BB fixes the first roll at 2. For AA to occur with BB, the second roll must be 2 (i.e., 2+2=42 + 2 = 4).

P(A∣B)=P(Second roll is 2)=P(2)=0.2P(A \mid B) = P(\text{Second roll is 2}) = P(2) = 0.2

This could also easily be examined from the graph the 2-2 is the only branch among the four possible cases under 2 in the first roll, so we have P(A∣B)=0.2×0.20.2=0.2P(A \mid B) = \frac{0.2\times 0.2}{0.2}=0.2.

Remark

We get P(A∣B)=0.2×0.20.2=0.2P(A \mid B) = \frac{0.2\times 0.2}{0.2}=0.2 by using the corollary to the numerator, which is the probability of getting a two twice. By the corollary, it can be expressed as the product of the probability we get 2 in the second trial given that 2 is ontained in the first attempt, and P(2)P(2). It is obvious that the condition does not work here, because we always have the same chance of getting a 2. Thus we have 0.220.2^2 as the numerator of P(A∣B)P(A \mid B).

2. P(B∣A)P(B \mid A)

The event AA (sum is 4) can occur through the following combinations:

  • (1,3)(1, 3) and (3,1)(3, 1)

  • (2,2)(2, 2)

Calculate P(A)P(A):

P(A)=P(1,3)+P(3,1)+P(2,2)=0.1⋅0.3+0.3⋅0.1+0.2⋅0.2=0.03+0.03+0.04=0.1P(A) = P(1, 3) + P(3, 1) + P(2, 2) = 0.1 \cdot 0.3 + 0.3 \cdot 0.1 + 0.2 \cdot 0.2 = 0.03 + 0.03 + 0.04 = 0.1
Remark

Note that:

  • We use Eq intasprob here, since for some event E,FE,F, P(E∩F)=P(E∣F)P(F)P(E\cap F) = P(E \mid F)P(F). We get the probability of each branch by using the RHS of the equation using the probability on the edge of the tree diagram. E.g., we have 3 for the second roll and 1 in the first roll, then we have 0.3×0.1=0.030.3\times 0.1 = 0.03 chance for the coresponding result.

  • The weight on the tree graph in every layer are treated as conditional probability, and 0.3 is the probability of a three given that the first roll gives 1

  • Also notice that all branches are disjoint, so we use additibity axiom here to sum up the probability.

Now, P(B∣A)P(B \mid A):

P(B∣A)=P(B∩A)P(A)=P(2,2)P(A)=0.040.1=0.4P(B \mid A) = \frac{P(B \cap A)}{P(A)} = \frac{P(2, 2)}{P(A)} = \frac{0.04}{0.1} = 0.4

From this example we see that despite that we have different probability for each branch of events, what we are doing to find conditional probability is still the same: subsetting the sample space and locate our target cases in that subspace.

We wrap up the use of lattice diagram and tree diagram here.

Lattice Diagram

Lattice diagrams are particularly useful for representing all possible outcomes in a sequence of events where outcomes are straightforward and often uniformly distributed, such as flipping coins or rolling dice. They are excellent for visualizing permutations, combinations, and the structure of outcomes in multi-stage experiments. A lattice diagram efficiently showcases how different paths can lead to various results, making it ideal for illustrating scenarios with equal probabilities or for modeling situations like financial derivatives pricing, where the path dependencies of options or other financial instruments are important.

Tree Diagram

Tree diagrams, on the other hand, excel in situations where outcomes have different probabilities(i.e., not uniformly distributed ) or the process is inherently hierarchical. They are invaluable for breaking down complex probability problems into manageable parts, especially in Bayesian statistics where updates to probability estimates are made as new information becomes available. Tree diagrams allow for the visualization of conditional probabilities and sequential decision processes, making them perfect for scenarios where events are dependent on previous outcomes, such as sequential games, Bayesian inference, or decision analysis.

But of course, we will not always use diagram for problem-solving, since you can imagine how complex the diagram could be when the scale of the problem increases. Fundamentally, we still need to obtain the result by analyzing the chain of events and their relations with the given conditions in different contexts.

Here are some more interesting examples.

Example

Celine is undecided whether to take a French course or a chemistry course. She estimates that her probability of receiving an A grade would be 12\frac{1}{2} in a French course and 23\frac{2}{3} in a chemistry course. If Celine decides to base her decision on the flip of a fair coin, what is the probability that she gets an A in chemistry?

Solution

Let CC be the event that Celine takes the chemistry course and AA be the event that she receives an A in whatever course she takes. The decision to take chemistry is based on the flip of a fair coin, thus P(C)=12P(C) = \frac{1}{2}.

Given CC (the event of taking chemistry), the probability of AA (receiving an A) in chemistry is P(A∣C)=23P(A \mid C) = \frac{2}{3}. Therefore, the probability that Celine takes chemistry and gets an A can be calculated using the rule of multiplication for probabilities:

P(CA)=P(C)P(A∣C)=(12)(23)=13.P(CA) = P(C)P(A \mid C) = \left(\frac{1}{2}\right)\left(\frac{2}{3}\right) = \frac{1}{3}.

Thus, the probability that Celine gets an A in chemistry, given that her decision is based on a coin flip, is 13\frac{1}{3}.

Remark

Some people may wonder why we are calculating probability of intersection instead of conditional probability here. Because semantically, the probability of getting A in chemistry given that we get head/tail of the coin seems equivalent to the probability of getting A in chemistry, and we get head/tail of the coin. The answer if completely NO. Because when we gauge the conditional probability, condition(s) matters. Will the result of flipping the coin affact the chance of getting A in chemistry? Of course no, so we are not finding a conditional probability. We will discuss this further in independence of events. This is a vivid example that demonstrates the difference of sumbolic language and natural language. The former outperforms in terms of accuracy but less understandable for human, while natural language is more understandable but in many cases, bring a lot of bias.

Example

Suppose that an urn contains 8 red balls and 4 white balls. We draw 2 balls from the urn without replacement.

(a) If we assume that at each draw, each ball in the urn is equally likely to be chosen, what is the probability that both balls drawn are red?

(b) Now suppose that the balls have different weights, with each red ball having weight rr and each white ball having weight ww. Suppose that the probability that a given ball in the urn is the next one selected is its weight divided by the sum of the weights of all balls currently in the urn. Now what is the probability that both balls are red?

Solution

(a) Let R1R_1 and R2R_2 denote, respectively, the events that the first and second balls drawn are red. Given that the first ball selected is red, there are 7 remaining red balls and 4 white balls, so P(R2∣R1)=711P(R_2|R_1) = \frac{7}{11}. As P(R1)P(R_1) is clearly 812\frac{8}{12}, the desired probability is calculated as:

P(R1R2)=P(R1)P(R2∣R1)=(23)(711)=1433.P(R_1 R_2) = P(R_1)P(R_2|R_1) = \left(\frac{2}{3}\right)\left(\frac{7}{11}\right) = \frac{14}{33}.

(b) For this part, we again let RiR_i be the event that the ii-th ball chosen is red and use:

P(R1R2)=P(R1)P(R2∣R1)P(R_1 R_2) = P(R_1)P(R_2|R_1)

Now, number the red balls, and let Bi,i=1,…,8B_i, i = 1, \ldots, 8 be the event that the first ball drawn is red ball number ii. Then:

P(R1)=P(⋃i=18Bi)=∑i=18P(Bi)=8r8r+4wP(R_1) = P\left(\bigcup_{i=1}^8 B_i\right) = \sum_{i=1}^8 P(B_i) = \frac{8r}{8r + 4w}

Moreover, given that the first ball is red, the urn then contains 7 red and 4 white balls. Thus, by a similar argument:

P(R2∣R1)=7r7r+4wP(R_2|R_1) = \frac{7r}{7r + 4w}

Hence, the probability that both balls are red is:

P(R1R2)=8r8r+4w×7r7r+4w=8r×7r(8r+4w)(7r+4w)P(R_1 R_2) = \frac{8r}{8r + 4w} \times \frac{7r}{7r + 4w} = \frac{8r \times 7r}{(8r + 4w)(7r + 4w)}

Now let's move further onto the generalization of corollary intersectascond, known as The multiplication rule of probability. Multiplication rule is used to calculate the probability of an event after a chain of event. In fact, we are using the previous definition multiple times here.

Theorem

If E1,E2,…,EnE_1, E_2, \ldots, E_n are events, then the probability of all these events occurring in sequence is given by:

P(E1E2⋯En)=P(E1)P(E2∣E1)P(E3∣E1E2)⋯P(En∣E1E2⋯En−1)P(E_1 E_2 \cdots E_n) = P(E_1) P(E_2 \mid E_1) P(E_3 \mid E_1 E_2) \cdots P(E_n \mid E_1 E_2 \cdots E_{n-1})
Proof

We prove the theorem by induction on the number of events nn.

Base case (n=2n = 2): For two events E1E_1 and E2E_2, the multiplication rule reduces to:

P(E1E2)=P(E1)P(E2∣E1),P(E_1 E_2) = P(E_1) P(E_2 \mid E_1),

which is the definition of the conditional probability of E2E_2 given E1E_1.

Inductive step: Assume that the theorem holds for n−1n-1 events. That is, we assume:

P(E1E2⋯En−1)=P(E1)P(E2∣E1)⋯P(En−1∣E1⋯En−2).P(E_1 E_2 \cdots E_{n-1}) = P(E_1) P(E_2 \mid E_1) \cdots P(E_{n-1} \mid E_1 \cdots E_{n-2}).

We need to prove that the rule holds for nn events. By the definition of conditional probability, we have:

P(E1E2⋯En)=P(E1E2⋯En−1)P(En∣E1E2⋯En−1).P(E_1 E_2 \cdots E_n) = P(E_1 E_2 \cdots E_{n-1}) P(E_n \mid E_1 E_2 \cdots E_{n-1}).

Applying the induction hypothesis, we substitute for P(E1E2⋯En−1)P(E_1 E_2 \cdots E_{n-1}):

P(E1E2⋯En)=(P(E1)P(E2∣E1)⋯P(En−1∣E1⋯En−2))P(En∣E1⋯En−1),P(E_1 E_2 \cdots E_n) = \left(P(E_1) P(E_2 \mid E_1) \cdots P(E_{n-1} \mid E_1 \cdots E_{n-2})\right) P(E_n \mid E_1 \cdots E_{n-1}),

which confirms the theorem for nn events.

Thus, by the principle of mathematical induction, the multiplication rule holds for any number of events.

Remark

Actually, we have a more decent way to prove it, which works like chain rule in differentiation.To prove the multiplication rule, just apply the definition of conditional probability to its right-hand side, giving

P(E1)P(E1E2)P(E1)P(E1E2E3)P(E1E2)⋯P(E1E2⋯En)P(E1E2⋯En−1)=P(E1E2⋯En).P(E_1)\frac{P(E_1E_2)}{P(E_1)}\frac{P(E_1E_2E_3)}{P(E_1E_2)}\cdots\frac{P(E_1E_2\cdots E_n)}{P(E_1E_2\cdots E_{n-1})}=P(E_1E_2\cdots E_n).
Example

An ordinary deck of 52 playing cards is randomly divided into 4 piles of 13 cards each. Compute the probability that each pile has exactly 1 ace.

Solution

Define events EiE_i, i=1,2,3,4i = 1, 2, 3, 4, as follows:

  • E1={the ace of spades is in any one of the piles}E_1 = \{\text{the ace of spades is in any one of the piles}\}

  • E2={the ace of spades and the ace of hearts are in different piles}E_2 = \{\text{the ace of spades and the ace of hearts are in different piles}\}

  • E3={the aces of spades, hearts, and diamonds are all in different piles}E_3 = \{\text{the aces of spades, hearts, and diamonds are all in different piles}\}

  • E4={all 4 aces are in different piles}E_4 = \{\text{all 4 aces are in different piles}\}

The desired probability is P(E1E2E3E4)P(E_1 E_2 E_3 E_4), and by the multiplication rule,

P(E1E2E3E4)=P(E1)P(E2∣E1)P(E3∣E1E2)P(E4∣E1E2E3)P(E_1 E_2 E_3 E_4) = P(E_1) P(E_2 \mid E_1) P(E_3 \mid E_1 E_2) P(E_4 \mid E_1 E_2 E_3)

Now, P(E1)=1P(E_1) = 1 since E1E_1 is the sample space SS. To determine P(E2∣E1)P(E_2 \mid E_1), consider the pile that contains the ace of spades. Because its remaining 12 cards are equally likely to be any 12 of the remaining 51 cards, the probability that the ace of hearts is among them is 12/5112/51, giving that

P(E2∣E1)=1−1251=3951P(E_2 \mid E_1) = 1 - \frac{12}{51} = \frac{39}{51}

Also, given that the ace of spades and ace of hearts are in different piles, it follows that the set of the remaining 24 cards of these two piles is equally likely to be any set of 24 of the remaining 50 cards. As the probability that the ace of diamonds is one of these 24 is 24/5024/50, we see that

P(E3∣E1E2)=1−2450=2650P(E_3 \mid E_1 E_2) = 1 - \frac{24}{50} = \frac{26}{50}

Because the same logic as used in the preceding yields that

P(E4∣E1E2E3)=1−3649=1349P(E_4 \mid E_1 E_2 E_3) = 1 - \frac{36}{49} = \frac{13}{49}

the probability that each pile has exactly 1 ace is

P(E1E2E3E4)=3951⋅2650⋅1349≈0.105P(E_1 E_2 E_3 E_4) = \frac{39}{51} \cdot \frac{26}{50} \cdot \frac{13}{49} \approx 0.105

That is, there is approximately a 10.5 percent chance that each pile will contain an ace. (Problem 13 gives another way of using the multiplication rule to solve this problem.)

Exercises

The locked source file leaves this exercise heading empty. Developed exercises resume after Bayes's theorem.

Bayes's Theorem

Bayes's Theorem and Bayesian Thinking

We can do more with conditional probability. Now consider event EE and FF. Regardless of what kind of events they are, we always have

E=EF∪EFc.E = EF \cup EF^c.

This makes sense not only from set theory's conclusion but also in terms of practicality. Because for any two events, assuming one of them will happen, then the other one could either happen or not happen. We know that EF and EFcEF \text{ and } EF^c are mutually exclusive, so by additivity axiom, we have

P(E)=P(EF)+P(EFc)=P(E∣F)P(F)+P(E∣Fc)P(Fc)=P(E∣F)P(F)+P(E∣Fc)[1−P(F)].\begin{aligned} P(E)& =P(EF) + P(EF^c) \\ &=P(E|F)P(F) + P(E|F^c)P(F^c) \\ &=P(E|F)P(F) + P(E|F^c)\bigl[1 - P(F)\bigr]. \end{aligned}

This can be interpreted as: the probability of the event E is measured by the weighted average of of the possibility that F happens and not happens. This formula enables us to get the probability of an event by conditioning the probability of some other events and its complement. This is very important for us, since some times we cannot get the probability of a certain event easily. The following example shows its advantage in finding conditional probability (in a reversed order).

Example

An insurance company believes that people can be divided into two classes: those who are accident prone and those who are not. The company's statistics show that an accident-prone person will have an accident at some time within a fixed 1-year period with probability .4, whereas this probability decreases to .2 for a person who is not accident prone. If we assume that 30 percent of the population is accident prone, what is the probability that a new policyholder will have an accident within a year of purchasing a policy?

Solution

We shall obtain the desired probability by first conditioning upon whether or not the policyholder is accident prone. Let A1A_1 denote the event that the policyholder will have an accident within a year of purchasing the policy, and let AA denote the event that the policyholder is accident prone. Hence, the desired probability is given by

P(A1)=P(A1∣A)P(A)+P(A1∣Ac)P(Ac)=(0.4)(0.3)+(0.2)(0.7)=0.26P(A_1) = P(A_1 \mid A)P(A) + P(A_1 \mid A^c)P(A^c) = (0.4)(0.3) + (0.2)(0.7) = 0.26

Note that in the previous example, we used the probability of A1A_1 given AA or AcA^c. The next example shows how we can get the probability of AA given A1A_1.

Example

Suppose that a new policyholder has an accident within a year of purchasing a policy. What is the probability that he or she is accident prone?

Solution

The desired probability is

P(A∣A1)=P(A∩A1)P(A1)=P(A1∣A)P(A)P(A1)=(0.3)(0.4)0.26=613P(A \mid A_1) = \frac{P(A \cap A_1)}{P(A_1)} = \frac{P(A_1 \mid A)P(A)}{P(A_1)} = \frac{(0.3)(0.4)}{0.26} = \frac{6}{13}

The figure uses an illustrative population of 1,0001{,}000 policyholders. Its event EE is the first-year accident event A1A_1 above. Filtering leaves 120+140=260120+140=260 people, so the accident-prone share becomes 120/260=6/13120/260=6/13.

Bayes filtering: 300 accident-prone and 700 other policyholders become 120 and 140 after conditioning on a first-year accident; the posterior is 6/13.

Bayes filtering: 300 accident-prone and 700 other policyholders become 120 and 140 after conditioning on a first-year accident; the posterior is 6/13.

With this example, we have determined that knowing the probability of an event EE (in this case, we can obtain its complement EcE^c using the complement rule such that P(Ec)=1−P(E)P(E^c) = 1 - P(E)), given another event FF and its complement FcF^c, enables us to calculate the probability of FF given EE.

This is known as Bayes's Theorem. The interesting part of the theorem lies not only in terms of algebra or practical significance, but it also illustrates a thinking pattern.

Theorem

Given two events AA and BB, with P(B)>0P(B) > 0, Bayes' theorem describes the probability of event AA given that event BB has occurred, and is formally defined as:

P(A∣B)=P(B∣A)P(A)P(B)P(A|B) = \frac{P(B|A)P(A)}{P(B)}

where:

  • P(A∣B)P(A|B) is the posterior probability of event AA occurring given the occurrence of event BB.

  • P(B∣A)P(B|A) is the likelihood, which is the probability of observing event BB given that event AA has occurred.

  • P(A)P(A) is the prior probability of event AA occurring.

  • P(B)P(B) is the marginal probability of event BB, which can also be calculated using the law of total probability if the complete set of outcomes related to AA is known:

P(B)=∑iP(B∣Ai)P(Ai)P(B) = \sum_{i} P(B|A_i)P(A_i)
Proof

Assume P(A)>0P(A)>0 as well as P(B)>0P(B)>0 so both displayed conditional probabilities are defined. The definition gives P(A∣B)P(B)=P(A∩B)=P(B∣A)P(A)P(A\mid B)P(B)=P(A\cap B)=P(B\mid A)P(A). Divide by P(B)P(B). In a total-probability denominator, the AiA_i must form a finite or countable disjoint partition; omit zero-probability cells before forming conditional terms. When P(A)=0P(A)=0, the posterior is directly zero by monotonicity, while P(B∣A)P(B\mid A) is undefined under this elementary event definition.

The Bayes's Formula is extremely concise and easy to understand, however profound philosophical and cognitive significance are combined into this expression.

Bayes's formula models the way human make judgment, of course, in a wise way. For each term of the formula, we have given an definition, which we need to figure out.

  • The prior probability is the existing belief in a certain event.

  • The likelihood is an factor that helps people make further judgment, which means the possibility of something else to happen under some given facts(in someone's mind, but may not a 100% true fact).

  • The marginal probability refers to the possibility of the other related event.

  • Posterior Probability refers to the new possibility given some other event(s) that fix/update the prior probability to a reasonable level.

If I tell you, this is exactly how human-beings , not just as an individual, but a civilization, learn, you may confuse that these seem very elusive and abstract, but we can clarify this by using one example.

Consider asking the question that "Is the earth flat or a sphere" to someone randomly picked from somewhere in the world, almost everyone will say yes, out of common sense. But clearly people do not always think so. In the past, or more specifically, before human-beings grasp the way of examine and measuring the nature, most people believe that we live on a flat land. However, with the advance of sailing technique, thrive of colonialism, and a cascade of geographical discovery, some people slightly change their mind, finding some tiny possibility that we live on a sphere. Only until the first global circling sail is done, that the earth is a sphere becomes acceptable to more and more people and finally become a common sense.

In this example, the prior probability of the event that the earth of is flat decreases, while the possibility of the contradictory event that the earth of sphere increases. We can analyze the numerator and denominator of the formula, and we will find that P(A)P(A) approaches to 0 and P(B)P(B) approaches to 1. Eventually, this makes the probability that earth is flat given all new known facts(observation from BB) to almost 0. This is exactly how we learn new things.

Though we know the interesting example, we still cannot prove it. We have shown how this makes sense previously, but only one thing is different here, since we consider that there could be any amount ii out comes for AA, while previously we used the base case where AA is a binary event, meaning it is only considered happen or not happen.

To work this problem out, we first introduces the law of total probability.

Theorem

Let B1,B2,…,BnB_1, B_2, \ldots, B_n be a finite or countably infinite partition of the sample space SS, such that Bi∩Bj=∅B_i \cap B_j = \emptyset for all i≠ji \neq j and ⋃i=1nBi=S\bigcup_{i=1}^n B_i = S, with P(Bi)>0P(B_i) > 0 for all ii. For any event AA in SS,

P(A)=∑i=1nP(A∣Bi)P(Bi).P(A) = \sum_{i=1}^n P(A \mid B_i) P(B_i).
Proof

Consider the event AA and the partition B1,B2,…,BnB_1, B_2, \ldots, B_n of the sample space SS. Since these sets are mutually exclusive and collectively exhaustive,

A=A∩S=A∩(⋃i=1nBi)=⋃i=1n(A∩Bi) (Set Distributive Law).A = A \cap S = A \cap \left( \bigcup_{i=1}^n B_i \right) = \bigcup_{i=1}^n (A \cap B_i) \text{ (Set Distributive Law)}.

Because the sets A∩BiA \cap B_i are mutually exclusive,

P(A)=P(⋃i=1n(A∩Bi))=∑i=1nP(A∩Bi).P(A) = P\left( \bigcup_{i=1}^n (A \cap B_i) \right) = \sum_{i=1}^n P(A \cap B_i).

Using the definition of conditional probability,

P(A∩Bi)=P(A∣Bi)P(Bi).P(A \cap B_i) = P(A \mid B_i) P(B_i).

Thus,

P(A)=∑i=1nP(A∣Bi)P(Bi),P(A) = \sum_{i=1}^n P(A \mid B_i) P(B_i),

which is the statement of the law of total probability.

Remark

You may also prove this by induction with n=2n=2 as a base case using the conclusion of the beginning of the section.

Now we can proceed to prove Bayes's Theorem.

Proof

Proof of Bayes's Theorem. Assume B1,B2,…,BnB_1, B_2, \ldots, B_n is a partition of the sample space, and AA is any event in the sample space. According to the law of total probability, we have:

P(A)=∑i=1nP(A∩Bi)P(A) = \sum_{i=1}^n P(A \cap B_i)

Using the definition of conditional probability, we express P(A∩Bi)P(A \cap B_i) as:

P(A∩Bi)=P(A∣Bi)P(Bi)P(A \cap B_i) = P(A | B_i)P(B_i)

Substituting this back into the law of total probability, we obtain:

P(A)=∑i=1nP(A∣Bi)P(Bi)P(A) = \sum_{i=1}^n P(A | B_i)P(B_i)

Now, consider the conditional probability P(Bj∣A)P(B_j | A) for some jj. By definition, it is:

P(Bj∣A)=P(A∩Bj)P(A)P(B_j | A) = \frac{P(A \cap B_j)}{P(A)}

Substituting P(A∩Bj)=P(A∣Bj)P(Bj)P(A \cap B_j) = P(A | B_j)P(B_j) into the equation, we get:

P(Bj∣A)=P(A∣Bj)P(Bj)P(A)P(B_j | A) = \frac{P(A | B_j)P(B_j)}{P(A)}

Since P(A)P(A) can be expanded using the law of total probability, it becomes:

P(Bj∣A)=P(A∣Bj)P(Bj)∑i=1nP(A∣Bi)P(Bi)P(B_j | A) = \frac{P(A | B_j)P(B_j)}{\sum_{i=1}^n P(A | B_i)P(B_i)}

This equation is Bayes' Theorem for the case where the partition {Bi}\{B_i\} consists of nn events. For the simpler case with only one BB and its complement BcB^c, Bayes' Theorem simplifies to:

P(B∣A)=P(A∣B)P(B)P(A∣B)P(B)+P(A∣Bc)P(Bc)P(B | A) = \frac{P(A | B)P(B)}{P(A | B)P(B) + P(A | B^c)P(B^c)}

Here are some typical problems.

Example

At a certain stage of a criminal investigation, the inspector in charge is 60%60\% convinced of the guilt of a certain suspect. Suppose, however, that a new piece of evidence which shows that the criminal has a certain characteristic (such as left-handedness, baldness, or brown hair) is uncovered. If 20%20\% of the population possesses this characteristic, how certain of the guilt of the suspect should the inspector now be if it turns out that the suspect has the characteristic?

Solution

Let GG denote the event that the suspect is guilty and CC the event that he possesses the characteristic of the criminal. We use Bayes' theorem to update our belief about GG given CC:

P(G∣C)=P(G∩C)P(C)P(G \mid C) = \frac{P(G \cap C)}{P(C)}

Applying the law of total probability to the denominator,

P(C)=P(C∣G)P(G)+P(C∣Gc)P(Gc)P(C) = P(C \mid G)P(G) + P(C \mid G^c)P(G^c)

Given that P(G)=0.60P(G) = 0.60 and P(Gc)=0.40P(G^c) = 0.40 (since the suspect being not guilty is the complement of the suspect being guilty), and assuming P(C∣G)=1.0P(C \mid G) = 1.0 (if the suspect is guilty, he definitely has the characteristic) and P(C∣Gc)=0.20P(C \mid G^c) = 0.20 (the probability that a non-guilty person has the characteristic),

P(C)=(1.0)(0.60)+(0.20)(0.40)=0.60+0.08=0.68P(C) = (1.0)(0.60) + (0.20)(0.40) = 0.60 + 0.08 = 0.68

Thus,

P(G∣C)=P(C∣G)P(G)P(C)=(1.0)(0.60)0.68≈0.882P(G \mid C) = \frac{P(C \mid G)P(G)}{P(C)} = \frac{(1.0)(0.60)}{0.68} \approx 0.882

This indicates that, given the suspect has the characteristic, the inspector should now be approximately 88.2%88.2\% certain of the suspect's guilt.

Example

Suppose that we have 3 cards that are identical in form, except that both sides of the first card are colored red, both sides of the second card are colored black, and one side of the third card is colored red and the other side black. The 3 cards are mixed up in a hat, and 1 card is randomly selected and put down on the ground. If the upper side of the chosen card is colored red, what is the probability that the other side is also colored red?

Solution

Let RRRR, BBBB, and RBRB denote the selected card, and let RR mean that its upper side is red. Assume each card is equally likely and either side is equally likely to face up. Then

P(R)=1⋅13+0⋅13+12⋅13=12.P(R)=1\cdot\frac13+0\cdot\frac13+\frac12\cdot\frac13=\frac12.

Given a red upper side, the lower side is red exactly when the card is RRRR. Since P(R∣RR)=1P(R\mid RR)=1,

P(RR∣R)=P(R∣RR)P(RR)P(R)=1⋅(1/3)1/2=23.P(RR\mid R)=\frac{P(R\mid RR)P(RR)}{P(R)}=\frac{1\cdot(1/3)}{1/2}=\frac23.

Equivalently, the three possible red upper faces consist of two faces from RRRR and one from RBRB. Two of these three outcomes have a red lower face. The remaining probability 1/31/3 is for a black lower face. The two possible card types are not equally likely after observing red.

Before applying conditional probability, specify how the observation is sampled. The next example depends on which child is selected to accompany the mother.

Example

A new couple, known to have two children, has just moved into town. Suppose that the mother is encountered walking with one of her children. If this child is a girl, what is the probability that both children are girls?

Solution

Assume the children's sexes are independent and equally likely, and that either child is equally likely to accompany the mother, independently of sex. Under this sampling model, define:

  • G1G_1: the first (that is, the oldest) child is a girl.

  • G2G_2: the second child is a girl.

  • GG: the child seen with the mother is a girl.

Also, let B1,B2,B_1, B_2, and BB denote similar events, except that "girl" is replaced by "boy." Now, the desired probability is P(G1G2∣G)P(G_1 G_2 \mid G), which can be expressed as follows:

P(G1G2∣G)=P(G1G2∩G)P(G)P(G_1 G_2 \mid G) = \frac{P(G_1 G_2 \cap G)}{P(G)}

where P(G)P(G) can be calculated using the law of total probability:

P(G)=P(G∣G1G2)P(G1G2)+P(G∣G1B2)P(G1B2)+P(G∣B1G2)P(B1G2)+P(G∣B1B2)P(B1B2)P(G) = P(G \mid G_1 G_2) P(G_1 G_2) + P(G \mid G_1 B_2) P(G_1 B_2) + P(G \mid B_1 G_2) P(B_1 G_2) + P(G \mid B_1 B_2) P(B_1 B_2)

Given that P(G∣G1G2)=1P(G \mid G_1 G_2) = 1 and P(G∣G1B2)=0.5P(G \mid G_1 B_2) = 0.5 and P(G∣B1G2)=0.5P(G \mid B_1 G_2) = 0.5 and P(G∣B1B2)=0P(G \mid B_1 B_2) = 0, and assuming that all four gender combinations are equally likely (14\frac{1}{4} each),

P(G)=1⋅14+12⋅14+12⋅14+0⋅14=12P(G) = 1 \cdot \frac{1}{4} + \frac{1}{2} \cdot \frac{1}{4} + \frac{1}{2} \cdot \frac{1}{4} + 0 \cdot \frac{1}{4} = \frac{1}{2}

Now, P(G1G2)P(G_1 G_2) is simply the probability of having two girls, which is 14\frac{1}{4}, so

P(G1G2∩G)=P(G1G2)=14P(G_1 G_2 \cap G) = P(G_1 G_2) = \frac{1}{4}

Therefore,

P(G1G2∣G)=1412=12P(G_1 G_2 \mid G) = \frac{\frac{1}{4}}{\frac{1}{2}} = \frac{1}{2}

Thus, if the child seen is a girl, the probability that both children are girls is 12\frac{1}{2}.

Independence of Events

Definition of Independence

previously, we defined the conditional probability of some event EE given FF by

P(E∣F)=P(EF)P(F).P(E\mid F)= \frac{P(EF)}{P(F)}.

You may recall that, in some examples earlier, we find P(E∣F)=P(E)P(E\mid F) = P(E), and we have P(E∣F)=P(E)  ⟺  P(EF)=P(E)×P(F)P(E\mid F)=P(E) \iff P(EF) = P(E) \times P(F). In this case, we claim that EE is independent of FF.

Definition

Two events E and F are said to be independent if and only if P(E∣F)=P(E)  ⟺  P(EF)=P(E)×P(F)P(E\mid F)=P(E) \iff P(EF) = P(E) \times P(F)

Here's a basic example.

Example

Suppose that we toss two fair dice. Consider the following events:

  • E1E_1: The event that the sum of the dice is 6.

  • FF: The event that the first die equals 4.

We are interested in determining whether these two events are independent.

Solution

First, calculate P(E1)P(E_1) and P(F)P(F), and then P(E1∩F)P(E_1 \cap F) to check for independence:

  • The probability that the sum of the dice equals 6, P(E1)P(E_1), can happen through the combinations (1,5),(2,4),(3,3),(4,2),(5,1)(1,5), (2,4), (3,3), (4,2), (5,1). Thus,
P(E1)=536P(E_1) = \frac{5}{36}

because there are 5 favorable outcomes out of 36 possible outcomes when two dice are thrown.

  • The probability that the first die equals 4, P(F)P(F), is simply,
P(F)=16P(F) = \frac{1}{6}

since one out of six faces of a die shows 4.

  • The probability of both E1E_1 and FF occurring, P(E1∩F)P(E_1 \cap F), happens only if the first die is 4 and the second die is 2 (to make the sum 6). Thus,
P(E1∩F)=136P(E_1 \cap F) = \frac{1}{36}

because there is only one favorable outcome for this combination under the condition that two dice are thrown.

Now, to check for independence, we examine if P(E1∩F)=P(E1)P(F)P(E_1 \cap F) = P(E_1)P(F):

P(E1)P(F)=(536)(16)=5216P(E_1)P(F) = \left(\frac{5}{36}\right)\left(\frac{1}{6}\right) = \frac{5}{216}

Clearly, 136≠5216\frac{1}{36} \neq \frac{5}{216}, indicating that E1E_1 and FF are not independent.

Another fact about independent event is that if two events are independent to each other, then they are also independent of each other's complement event.

Proposition

If EE and FF are independent, then so are EE and FcF^c.

Proof

Assume that EE and FF are independent. Since EE can be expressed as the union of disjoint events EFEF and EFcEF^c, we write

P(E)=P(EF)+P(EFc).P(E) = P(EF) + P(EF^c).

Using the independence of EE and FF, we have

P(EF)=P(E)P(F).P(EF) = P(E)P(F).

Thus, the probability of EE can be rewritten using the complement rule as

P(E)=P(E)P(F)+P(EFc).P(E) = P(E)P(F) + P(EF^c).

Rearranging terms gives

P(EFc)=P(E)−P(E)P(F)=P(E)(1−P(F))=P(E)P(Fc).P(EF^c) = P(E) - P(E)P(F) = P(E)(1 - P(F)) = P(E)P(F^c).

Hence, P(EFc)=P(E)P(Fc)P(EF^c) = P(E)P(F^c) shows that EE and FcF^c are independent.

Multiple Independence

We have discussed the base case of independence between two events, how it is like for more events? Here is an example.

Example

Two fair dice are thrown. Let EE denote the event that the sum of the dice is 7. Let FF denote the event that the first die equals 4 and GG denote the event that the second die equals 3. From earlier discussions, we know that EE is independent of FF, and the same reasoning shows that EE is also independent of GG. However, we find that EE is not independent of the joint event FGFG because P(E∣FG)=1P(E \mid FG) = 1. Given:

  • EE: Sum of two dice is 7.

  • FF: First die is 4.

  • GG: Second die is 3.

Solution

Independence Analysis:

  • EE and FF are independent:
P(E∩F)=P(E)P(F)  ⟹  136=(16)(16)=136P(E \cap F) = P(E)P(F) \implies \frac{1}{36} = \left(\frac{1}{6}\right)\left(\frac{1}{6}\right) = \frac{1}{36}
  • EE and GG are independent:
P(E∩G)=P(E)P(G)  ⟹  136=(16)(16)=136P(E \cap G) = P(E)P(G) \implies \frac{1}{36} = \left(\frac{1}{6}\right)\left(\frac{1}{6}\right) = \frac{1}{36}
  • EE and FGFG are not independent. If FF and GG both occur, the sum is automatically 7, so
P(E∣FG)=1.P(E \mid FG) = 1.

Then

P(E∩FG)=P(FG)≠P(E)P(FG).P(E \cap FG) = P(FG) \neq P(E)P(FG).

Indeed, P(FG)=136P(FG) = \frac{1}{36} and P(E)=16P(E) = \frac{1}{6}, so

P(E)P(FG)=(16)(136)=1216≠136.P(E)P(FG) = \left(\frac{1}{6}\right)\left(\frac{1}{36}\right) = \frac{1}{216} \neq \frac{1}{36}.

Thus, while EE is independent of both FF and GG individually, it is not independent of the joint event FGFG, illustrating how independence between individual events does not necessarily extend to independence with joint events.

Definition

Three events EE, FF, and GG are said to be independent if the following conditions hold:

P(EFG)=P(E)P(F)P(G),P(EF)=P(E)P(F),P(EG)=P(E)P(G),P(FG)=P(F)P(G).\begin{aligned} P(EFG) &= P(E)P(F)P(G), \\ P(EF) &= P(E)P(F), \\ P(EG) &= P(E)P(G), \\ P(FG) &= P(F)P(G). \end{aligned}

Note that if E,FE,F, and GG are independent, then EE will be independent of any event formed from FF and G.G. For instance, EE is independent of F∪GF\cup G,since

P[E(F∪G)]=P(EF∪EG)=P(EF)+P(EG)−P(EFG)=P(E)P(F)+P(E)P(G)−P(E)P(FG)=P(E)[P(F)+P(G)−P(FG)]=P(E)P(F∪G)\begin{aligned}P[E(F\cup G)]&=P(EF\cup EG)\\&=P(EF)+P(EG)-P(EFG)\\&=P(E)P(F)+P(E)P(G)-P(E)P(FG)\\&=P(E)[P(F)+P(G)-P(FG)]\\&=P(E)P(F\cup G)\end{aligned}

You may find that this notion is kind of like inclusion-exclusion, and the only difference is that we are not taking the case of one object into account. We can easily extend the definition of independence to infinitely many events, which will be an exercise probelm.

Example

Consider an infinite sequence of independent trials, each resulting in a success with probability pp and a failure with probability 1−p1 - p. Determine the probability that:

  1. At least 1 success occurs in the first nn trials.

  2. Exactly kk successes occur in the first nn trials.

  3. All trials result in successes.

Solution

(a) Probability of at least one success in the first nn trials

To find this probability, it's easiest to first compute the probability of the complementary event: no successes in the first nn trials. If EiE_i denotes a failure on the ii-th trial, by independence, the probability of no successes is:

P(E1E2…En)=P(E1)P(E2)⋯P(En)=(1−p)nP(E_1 E_2 \ldots E_n) = P(E_1)P(E_2) \cdots P(E_n) = (1-p)^n

Thus, the probability of at least one success is:

P(at least one success)=1−P(no successes)=1−(1−p)nP(\text{at least one success}) = 1 - P(\text{no successes}) = 1 - (1-p)^n

(b) Probability of exactly kk successes in the first nn trials

Consider any specific sequence of nn outcomes containing kk successes and n−kn - k failures. Each sequence occurs with probability pk(1−p)n−kp^k (1-p)^{n-k}, and the number of such sequences is given by the binomial coefficient (nk)\binom{n}{k}:

P(exactly k successes)=(nk)pk(1−p)n−kP(\text{exactly } k \text{ successes}) = \binom{n}{k} p^k (1-p)^{n-k}

(c) Probability that all trials result in successes By part (a), the probability that the first nn trials all result in success is pnp^n. The event that all trials result in success is the intersection of the events of success on each trial, EiE_i, over an infinite number of trials. Using the continuity property of probabilities:

P(⋂i=1∞Ei)=lim⁡n→∞P(E1E2…En)=lim⁡n→∞pnP(\bigcap_{i=1}^\infty E_i) = \lim_{n \to \infty} P(E_1 E_2 \ldots E_n) = \lim_{n \to \infty} p^n

This limit is 0 if p<1p < 1 and 1 if p=1p = 1, reflecting that only if success is certain on every trial (i.e., p=1p = 1) will we surely have success on every trial in an infinite sequence.

Example

Consider independent trials consisting of rolling a pair of fair dice. What is the probability that an outcome of 5 appears before an outcome of 7 when the outcome of a roll is the sum of the dice?

Solution

If we let EnE_n denote the event that no 5 or 7 appears on the first n−1n-1 trials and a 5 appears on the nn-th trial, then the desired probability is:

P(⋃n=1∞En)=∑n=1∞P(En)P\left( \bigcup_{n=1}^\infty E_n \right) = \sum_{n=1}^\infty P(E_n)

Given P(5 on any trial)=436P(5 \text{ on any trial}) = \frac{4}{36} and P(7 on any trial)=636P(7 \text{ on any trial}) = \frac{6}{36}, by the independence of trials, we compute:

P(En)=(1−1036)n−1×436P(E_n) = \left(1 - \frac{10}{36}\right)^{n-1} \times \frac{4}{36}

Therefore, the probability is:

P(⋃n=1∞En)=19∑n=1∞(1318)n−1=19×11−1318=25P\left( \bigcup_{n=1}^\infty E_n \right) = \frac{1}{9} \sum_{n=1}^\infty \left(\frac{13}{18}\right)^{n-1} = \frac{1}{9} \times \frac{1}{1 - \frac{13}{18}} = \frac{2}{5}

Alternative Method: The probability can also be derived using conditional probabilities:

P(E)=P(E∣F)P(F)+P(E∣G)P(G)+P(E∣H)P(H)P(E∣F)=1,P(E∣G)=0,P(E∣H)=P(E)P(E)=19+1318P(E)\begin{aligned} P(E) &= P(E|F)P(F) + P(E|G)P(G) + P(E|H)P(H) \\ P(E|F) &= 1, \quad P(E|G) = 0, \quad P(E|H) = P(E) \\ P(E) &= \frac{1}{9} + \frac{13}{18} P(E) \end{aligned}

Solving for P(E)P(E), we get P(E)=25P(E) = \frac{2}{5}.

Remark

The answer is intuitive as the probability of a 5 occurring on any roll is 436\frac{4}{36} and for a 7 is 636\frac{6}{36}. Thus, the probability that a 5 appears before a 7 should be 410\frac{4}{10}, as indeed it is.

if EE and FF are mutually exclusive events of an experiment, then, when independent trials of the experiment are performed, the event EE will occur before the event FF with probability

P(E)P(E)+P(F).\frac{P(E)}{P(E)+P(F)}.

Exercises

Exercise

Suppose rolling a fair dice twice, is the event that the sum of two trial is 7 independent of the first roll? Prove or disprove it. Explain the reason.

Proof

Let AA represent the event that the sum of the dice is 7, and BiB_i the event that the first die shows ii. The events are independent if P(A∩Bi)=P(A)P(Bi)P(A \cap B_i) = P(A)P(B_i) for each ii.

Calculation:

  • The probability that the sum equals 7, P(A)P(A), is the number of favorable outcomes (1,6),(2,5),(3,4),(4,3),(5,2),(6,1)(1,6), (2,5), (3,4), (4,3), (5,2), (6,1) over the total outcomes when rolling two dice:
P(A)=636=16.P(A) = \frac{6}{36} = \frac{1}{6}.
  • The probability of rolling any specific number on a fair die, P(Bi)P(B_i), is:
P(Bi)=16.P(B_i) = \frac{1}{6}.
  • The probability of both AA and BiB_i occurring simultaneously, P(A∩Bi)P(A \cap B_i), happens only when the first die is ii and the second die rolls 7−i7-i. Thus:
P(A∩Bi)=136.P(A \cap B_i) = \frac{1}{36}.

Checking Independence: Since the product of P(A)P(A) and P(Bi)P(B_i) is:

P(A)P(Bi)=16×16=136,P(A)P(B_i) = \frac{1}{6} \times \frac{1}{6} = \frac{1}{36},

and this is equal to P(A∩Bi)P(A \cap B_i), the equality:

P(A∩Bi)=P(A)P(Bi)P(A \cap B_i) = P(A)P(B_i)

holds for each ii, proving that AA and BiB_i are independent. Knowing the outcome of the first die roll does not affect the probability of the sum being 7.

This conclusion follows from the fact that the condition required for AA depends equally on both dice, and the outcome of one does not skew the likelihood of achieving a total of 7 compared to any other total.

Exercise

Use the definition below to show that each event is independent of every Boolean combination of the remaining events.

Theorem

A set of events {E1,E2,…,En}\{E_1, E_2, \dots, E_n\} is said to be mutually independent if and only if for every nonempty subset {Ei1,Ei2,…,Eik}\{E_{i_1}, E_{i_2}, \dots, E_{i_k}\} of these events, the probability of the intersection of these events equals the product of their probabilities:

P(⋂j=1kEij)=∏j=1kP(Eij).P\left(\bigcap_{j=1}^k E_{i_j}\right) = \prod_{j=1}^k P(E_{i_j}).

This must hold for every kk with 1≤k≤n1 \leq k \leq n and for every subset of indices i1,i2,…,iki_1, i_2, \dots, i_k.

Proof

Assume {E1,E2,…,En}\{E_1, E_2, \dots, E_n\} is a set of mutually independent events. We need to show that for any ii and for any event AA formed from {E1,…,Ei−1,Ei+1,…,En}\{E_1, \dots, E_{i-1}, E_{i+1}, \dots, E_n\} using any Boolean operations, the events EiE_i and AA are independent, i.e., P(Ei∩A)=P(Ei)P(A)P(E_i \cap A) = P(E_i)P(A).

Since AA can be expressed as a union of intersections of events and their complements from {E1,…,Ei−1,Ei+1,…,En}\{E_1, \dots, E_{i-1}, E_{i+1}, \dots, E_n\}, we apply the principle of inclusion-exclusion and the mutual independence assumption:

P(A)=P(⋃(Ej or Ejc))=sum of P(⋂(Ej or Ejc)) terms.P(A) = P\left(\bigcup (E_j \text{ or } E_j^c)\right) = \text{sum of } P\left(\bigcap (E_j \text{ or } E_j^c)\right) \text{ terms}.

Each term P(⋂(Ej or Ejc))P\left(\bigcap (E_j \text{ or } E_j^c)\right) reduces to the product of probabilities of EjE_j or 1−P(Ej)1 - P(E_j) by independence.

For Ei∩AE_i \cap A, apply the same decomposition:

P(Ei∩A)=P(Ei∩⋃(Ej or Ejc))=sum of P(Ei∩⋂(Ej or Ejc)) terms.P(E_i \cap A) = P\left(E_i \cap \bigcup (E_j \text{ or } E_j^c)\right) = \text{sum of } P(E_i \cap \bigcap (E_j \text{ or } E_j^c)) \text{ terms}.

By mutual independence, each intersection involving EiE_i simplifies to P(Ei)P(E_i) times the product of probabilities from the other events or their complements, verifying that:

P(Ei∩A)=P(Ei)P(A).P(E_i \cap A) = P(E_i)P(A).

Thus, EiE_i is independent of any Boolean combination of the remaining events, as required.

Further Conditional Probability

Probability Axiom in Conditional Probability

In the previous sections we introduced conditional probability and Bayes's formula. The next step is to push those identities further so that more problems can be solved.

We introduced the three axioms of probability and many propositions derived from them. While conditional probability is also a kind of "special" probability, these rules should also work for them, but as usual, we must provide rigorous proof.

Proposition

Let EE and FF be events, and let EiE_i, for i=1,2,…i=1, 2, \ldots, be a sequence of mutually exclusive events. Then:

  • 0≤P(E∣F)≤10 \leq P(E \mid F) \leq 1.

  • P(S∣F)=1P(S \mid F) = 1 where SS is the sample space.

  • If EiE_i, i=1,2,…i = 1, 2, \ldots, are mutually exclusive, then

P(⋃i=1∞Ei∣F)=∑i=1∞P(Ei∣F).P\left(\bigcup_{i=1}^\infty E_i \mid F\right) = \sum_{i=1}^\infty P(E_i \mid F).
Proof

To prove part (a), note that since E∩F⊆FE \cap F \subseteq F, it follows that P(E∩F)≤P(F)P(E \cap F) \leq P(F), hence 0≤P(E∣F)≤10 \leq P(E \mid F) \leq 1 as P(E∣F)=P(E∩F)P(F)P(E \mid F) = \frac{P(E \cap F)}{P(F)}. For part (b), since S∩F=FS \cap F = F, we have P(S∣F)=P(F)P(F)=1P(S \mid F) = \frac{P(F)}{P(F)} = 1. Part (c) follows from the properties of probability measures over countable unions of disjoint sets and the definition of conditional probability:

P(⋃i=1∞Ei∣F)=P(⋃i=1∞(Ei∩F))P(F)=∑i=1∞P(Ei∩F)P(F)=∑i=1∞P(Ei∣F).P\left(\bigcup_{i=1}^\infty E_i \mid F\right) = \frac{P\left(\bigcup_{i=1}^\infty (E_i \cap F)\right)}{P(F)} = \frac{\sum_{i=1}^\infty P(E_i \cap F)}{P(F)} = \sum_{i=1}^\infty P(E_i \mid F).

Conditioning therefore preserves the probability axioms.

Proposition

If we define Q(E)=P(E∣F)Q(E) = P(E \mid F), then QQ can be regarded as a probability function on the events of SS, and the propositions previously proved for probabilities apply to Q(E)Q(E). For instance,

Q(E1∪E2)=Q(E1)+Q(E2)−Q(E1∩E2).Q(E_1 \cup E_2) = Q(E_1) + Q(E_2) - Q(E_1 \cap E_2).
Proof

Fix P(F)>0P(F)>0. Nonnegativity, normalization, and countable additivity were proved immediately above, which are exactly the probability axioms for QQ. For the displayed identity, multiply ordinary inclusion-exclusion for E1∩FE_1\cap F and E2∩FE_2\cap F by 1/P(F)1/P(F). Their intersection is E1∩E2∩FE_1\cap E_2\cap F, producing the stated conditional formula.

Remark

It means that all conclusions about probability that we have proven are applicable to conditional probability. Also, do check exercise 1 of this section before moving on to the next example.

Example

Consider a scenario involving an insurance company that categorizes new policyholders into two groups: those who are accident prone and those who are not. It is known that:

  • The probability that an accident-prone person has an accident in any given year is 0.4.

  • The probability that a person who is not accident-prone has an accident in any given year is 0.2.

  • The proportion of the accident-prone population among new policyholders is 310\frac{3}{10}.

Given that a new policyholder has had an accident in the first year of their policy, what is the probability that they will have another accident in the second year?

Solution

Let AA denote the event that the policyholder is accident-prone, and let A1A_1 and A2A_2 denote accidents in the first and second years. Assume that the policyholder's type stays fixed and that yearly accidents are conditionally independent given that type. The annual probabilities alone would not determine the answer without a dependence assumption.

From the first-year accident, Bayes' rule gives

P(A∣A1)=(0.4)(0.3)0.26=613,P(Ac∣A1)=713.P(A\mid A_1)=\frac{(0.4)(0.3)}{0.26}=\frac{6}{13}, \qquad P(A^c\mid A_1)=\frac{7}{13}.

Conditional independence gives P(A2∣A∩A1)=0.4P(A_2\mid A\cap A_1)=0.4 and P(A2∣Ac∩A1)=0.2P(A_2\mid A^c\cap A_1)=0.2. Therefore,

P(A2∣A1)=P(A2∣A∩A1)P(A∣A1)+P(A2∣Ac∩A1)P(Ac∣A1)=25613+15713=1965≈0.2923.\begin{aligned} P(A_2\mid A_1) &=P(A_2\mid A\cap A_1)P(A\mid A_1) +P(A_2\mid A^c\cap A_1)P(A^c\mid A_1)\\ &=\frac{2}{5}\frac{6}{13}+\frac{1}{5}\frac{7}{13}\\ &=\frac{19}{65}\approx 0.2923. \end{aligned}
ExampleUpdating a posterior with successive observations

Let H1,…,HnH_1,\ldots,H_n be mutually exclusive, exhaustive hypotheses. Assume the observed events have positive probability, and omit any zero-posterior hypotheses from conditional expressions.

After observing E1E_1,

P(Hi∣E1)=P(E1∣Hi)P(Hi)P(E1).P(H_i\mid E_1)=\frac{P(E_1\mid H_i)P(H_i)}{P(E_1)}.

The general update after observing E2E_2 retains E1E_1 in the new likelihood:

P(Hi∣E1,E2)=P(E2∣Hi,E1)P(Hi∣E1)P(E2∣E1),P(H_i\mid E_1,E_2) =\frac{P(E_2\mid H_i,E_1)P(H_i\mid E_1)}{P(E_2\mid E_1)},

where

P(E2∣E1)=∑jP(E2∣Hj,E1)P(Hj∣E1).P(E_2\mid E_1)=\sum_j P(E_2\mid H_j,E_1)P(H_j\mid E_1).

If E1E_1 and E2E_2 are conditionally independent given each hypothesis, the likelihood simplifies to P(E2∣Hi)P(E_2\mid H_i). Then

P(Hi∣E1,E2)=P(E2∣Hi)P(E1∣Hi)P(Hi)P(E2∣E1)P(E1).P(H_i\mid E_1,E_2) =\frac{P(E_2\mid H_i)P(E_1\mid H_i)P(H_i)}{P(E_2\mid E_1)P(E_1)}.

The denominator is the joint evidence P(E1∩E2)P(E_1\cap E_2), a product rather than a ratio. For example, choose one coin of unknown identity from a collection with known biases, then flip that same coin repeatedly. Conditional on its identity, independent flips let us update using the current posterior and the next likelihood, without storing every earlier outcome separately.

Multi-conditional Probability

We can also generalize conditional event with multiple conditions. All we need to do is taking the intersection of the rest of the conditions(except one of the condition), as a single event. This could be proven easily with mathematical induction.

Definition

The multi-conditional probability of an event AA given multiple events B1,B2,…,BnB_1, B_2, \ldots, B_n is defined by the formula:

P(A∣B1,B2,…,Bn)=P(A∩B1∩B2∩…∩Bn)P(B1∩B2∩…∩Bn)P(A \mid B_1, B_2, \ldots, B_n) = \frac{P(A \cap B_1 \cap B_2 \cap \ldots \cap B_n)}{P(B_1 \cap B_2 \cap \ldots \cap B_n)}

provided that P(B1∩B2∩…∩Bn)>0P(B_1 \cap B_2 \cap \ldots \cap B_n) > 0.

RemarkMultiple conditions form one event

Set B=B1∩⋯∩BnB=B_1\cap\cdots\cap B_n. When P(B)>0P(B)>0, ordinary conditional probability gives

P(A∣B1,…,Bn)=P(A∣B)=P(A∩B)P(B).P(A\mid B_1,\ldots,B_n) =P(A\mid B) =\frac{P(A\cap B)}{P(B)}.

This is an application of the definition to an intersection, so no induction argument is needed. The events BiB_i need not be independent.

Example

A computer system's reliability is critical to an organization. The failure of this system can be influenced by multiple factors, some of which are interdependent. These factors include software malfunction (S), hardware malfunction (H), network failure (N), power surges (P), and user error (U).

  • P(Failure∣S∩H∩N∩P∩U)=0.90P(\text{Failure} | S \cap H \cap N \cap P \cap U) = 0.90

  • P(S∩H∩N∩P∩U)=0.01P(S \cap H \cap N \cap P \cap U) = 0.01

  • P(Failure∣S∩H∩N)=0.75P(\text{Failure} | S \cap H \cap N) = 0.75

  • P(S∩H∩N)=0.05P(S \cap H \cap N) = 0.05

  • P(P∩U∣S∩H∩N)=0.20P(P \cap U | S \cap H \cap N) = 0.20

    (Conditional probability reflecting the likelihood of power surges and user error given the first three conditions)

Assess the comprehensive probability of system failure incorporating all these factors, taking into account their conditional dependencies.

Solution

First, calculate the joint probability of all factors:

P(S∩H∩N∩P∩U)=P(P∩U∣S∩H∩N)×P(S∩H∩N)P(S \cap H \cap N \cap P \cap U) = P(P \cap U | S \cap H \cap N) \times P(S \cap H \cap N)P(S∩H∩N∩P∩U)=0.20×0.05=0.01P(S \cap H \cap N \cap P \cap U) = 0.20 \times 0.05 = 0.01

Then, apply the multi-conditional probability:

P(Failure∣S,H,N,P,U)=P(Failure∩S∩H∩N∩P∩U)P(S∩H∩N∩P∩U)P(\text{Failure} | S, H, N, P, U) = \frac{P(\text{Failure} \cap S \cap H \cap N \cap P \cap U)}{P(S \cap H \cap N \cap P \cap U)}P(Failure∩S∩H∩N∩P∩U)=P(Failure∣S∩H∩N∩P∩U)×P(S∩H∩N∩P∩U)P(\text{Failure} \cap S \cap H \cap N \cap P \cap U) = P(\text{Failure} | S \cap H \cap N \cap P \cap U) \times P(S \cap H \cap N \cap P \cap U)P(Failure∩S∩H∩N∩P∩U)=0.90×0.01=0.009P(\text{Failure} \cap S \cap H \cap N \cap P \cap U) = 0.90 \times 0.01 = 0.009

Therefore, the comprehensive probability of system failure given all factors are:

P(Failure∣S,H,N,P,U)=0.0090.01=0.90P(\text{Failure} | S, H, N, P, U) = \frac{0.009}{0.01} = 0.90

Exercises

Exercise

Show that proposition ex

holds and thus prove that

Q(E1∪E2)=Q(E1)+Q(E2)−Q(E1E2)Q(E_1 \cup E_2)=Q(E_1) + Q(E_2) - Q(E_1E_2)  ⟺  P(E1∣F)=P(E1∣E2F)P(E2∣F)+P(E1∣E2cF)P(E2c∣F)\iff P(E_1|F)=P(E_1|E_2F)P(E_2|F) + P(E_1|E_2^cF)P(E_2^c|F)
Proof

To show that QQ is a probability measure, we must verify that it satisfies the axioms of probability. Since Q(E)=P(E∣F)Q(E) = P(E \mid F) and PP is a probability measure:

  1. Q(S)=P(S∣F)=1Q(S) = P(S \mid F) = 1, since S∩F=FS \cap F = F.

  2. Q(E)≥0Q(E) \geq 0 for any event EE, as conditional probabilities are non-negative.

  3. For any sequence of mutually exclusive events E1,E2,…E_1, E_2, \ldots,

Q(⋃i=1∞Ei)=P(⋃i=1∞Ei∣F)=∑i=1∞P(Ei∣F)=∑i=1∞Q(Ei),Q\left(\bigcup_{i=1}^\infty E_i\right) = P\left(\bigcup_{i=1}^\infty E_i \mid F\right) = \sum_{i=1}^\infty P(E_i \mid F) = \sum_{i=1}^\infty Q(E_i),

by the sigma additivity of PP and the definition of conditional probability.

Thus, QQ behaves as a probability measure on SS, and by the properties of probability measures, the formula for the union of two events follows.

For the second part, assume Q(E2)>0Q(E_2)>0 and Q(E2c)>0Q(E_2^c)>0. The conditional probability under QQ is

Q(E1∣E2)=Q(E1∩E2)Q(E2).Q(E_1\mid E_2)=\frac{Q(E_1\cap E_2)}{Q(E_2)}.

We can get the total probability of E1E_1 by

Q(E1)=Q(E1∣E2)Q(E2)+Q(E1∣E2c)Q(E2c).Q(E_{1})=Q(E_{1}|E_{2})Q(E_{2}) + Q(E_{1}|E_{2}^{c})Q(E_{2}^{c}).

We also have

Q(E1∣E2)=Q(E1E2)Q(E2)=P(E1E2∣F)P(E2∣F)=P(E1E2F)P(F)P(E2F)P(F)=P(E1∣E2F).Q(E_1|E_2) = \frac{Q(E_1E_2)}{Q(E_2)} =\frac{P(E_1E_2|F)}{P(E_2|F)} =\frac{\frac{P(E_1E_2F)}{P(F)}}{\frac{P(E_2F)}{P(F)}} =P(E_1|E_2F).

Similarly, Q(E1∣E2c)=P(E1∣E2cF)Q(E_1|E_2^c) = P(E_1|E_2^cF).

Substituting, we have

P(E1∣F)=P(E1∣E2F)P(E2∣F)+P(E1∣E2cF)P(E2c∣F).P(E_1|F)=P(E_1|E_2F)P(E_2|F) + P(E_1|E_2^cF)P(E_2^c|F).

If a branch has zero probability under QQ, omit it; no conditional probability on a null event is needed. This completes the proof.