Skip to content
VidiMaster it, module by module
Module 6/Conditional Probability & Independence

Bayes' Formula: Reversing the Order of Conditioning

Conditional probability answers P(A∣B)P(A\mid B); Bayes' formula (Proposition 11.2) answers the reverse question P(B∣A)P(B\mid A) -- the probability of a cause given an observed effect. Assembled from the definition of conditional probability (Definition 11.1) and the law of total probability (Proposition 11.1), it updates a prior belief P(B)P(B) into a posterior P(B∣A)P(B\mid A) once evidence AA arrives (Remark 11.2). The centerpiece is the medical-test / base-rate problem (Example 11.6), where a highly accurate test for a rare disease returns a surprisingly low posterior -- the base-rate fallacy that routinely fools intuition.

Before you start — give these a try

Attempting first primes your brain for the lesson — even if you miss. Nothing is graded or saved; it's just a warm-up.

A company runs two production lines. Line A makes 60%60\% of the units and has a 2%2\% defect rate; line B makes the other 40%40\% and has a 5%5\% defect rate. A unit is pulled from the warehouse and found to be defective. Find P(line A∣defective)P(\text{line A}\mid\text{defective}), to three decimal places.

In the medical-test setting, the prior probability of disease is P(D)=0.005P(D)=0.005, and a positive result gives the posterior P(D∣+)≈0.194P(D\mid +)\approx0.194 (Example 11.6). Compared with the prior, the posterior is:

What you’ll be able to do

  • State Proposition 11.2 (Bayes' Formula): for P(A),P(B),P(Bc)>0P(A),P(B),P(B^c)>0, P(B∣A)=P(A∣B)P(B)P(A∣B)P(B)+P(A∣Bc)P(Bc)P(B\mid A)=\dfrac{P(A\mid B)P(B)}{P(A\mid B)P(B)+P(A\mid B^c)P(B^c)}, and use it to reverse the order of conditioning.
  • Derive Bayes' formula from Definition 11.1 (conditional probability, via the multiplication rule) together with Proposition 11.1 (the law of total probability), which expands the denominator into P(A)P(A).
  • Distinguish the prior P(B)P(B) from the posterior P(B∣A)P(B\mid A) (Remark 11.2), and explain how observing AA updates belief depending on how diagnostic AA is.
  • Solve the medical-test base-rate problem (Example 11.6) from a test's sensitivity P(+∣D)P(+\mid D), its false-positive rate P(+∣Dc)P(+\mid D^c) (specificity 1−P(+∣Dc)1-P(+\mid D^c)), and the prevalence P(D)P(D).
  • Recognize the base-rate fallacy -- why high sensitivity does not make a positive result a near-certainty for a rare condition -- and decide whether a posterior is larger or smaller than its prior.

In your course

· MATH2015 · Linear Algebra & Probability
§11.2 Bayes' Formula§11.1 Conditional Probability
  • Proposition 11.2Bayes' Formula
    For P(A),P(B),P(Bc)>0P(A),P(B),P(B^c)>0, P(B∣A)=P(A∣B)P(B)P(A∣B)P(B)+P(A∣Bc)P(Bc)P(B\mid A)=\dfrac{P(A\mid B)P(B)}{P(A\mid B)P(B)+P(A\mid B^c)P(B^c)}.
  • Remark 11.2Prior and posterior probability
    P(B)P(B) is the prior (belief before evidence); P(B∣A)P(B\mid A) is the posterior (belief after observing AA).
  • Example 11.6Medical test (base-rate problem)
    Sensitivity 0.960.96, false-positive rate 0.020.02, prevalence 0.0050.005 give P(D∣+)=48247≈0.194P(D\mid +)=\dfrac{48}{247}\approx0.194.
  • Example 11.5Two urns -- reversing conditioning
    P(urn I∣red)=511≈0.455<12P(\text{urn I}\mid\text{red})=\dfrac{5}{11}\approx0.455<\tfrac12, so a red ball is more likely from urn II.
  • Proposition 11.1Law of Total Probability
    For a partition {Bi}\{B_i\}, P(A)=∑iP(A∣Bi)P(Bi)P(A)=\sum_i P(A\mid B_i)P(B_i); two-event form P(A)=P(A∣B)P(B)+P(A∣Bc)P(Bc)P(A)=P(A\mid B)P(B)+P(A\mid B^c)P(B^c).
  • Definition 11.1Conditional Probability
    P(A∣B)=P(A∩B)P(B)P(A\mid B)=\dfrac{P(A\cap B)}{P(B)} for P(B)>0P(B)>0; multiplication rule P(A∩B)=P(A∣B)P(B)P(A\cap B)=P(A\mid B)P(B).
Grounded in MATH2015 §11.2 (Bayes' Formula), with conditional-probability and law-of-total-probability background from §11.1. Every posterior was verified via Bayes' formula with the law of total probability in the denominator.
1

Reversing the order of conditioning

Conditional probability (Definition 11.1) gives P(A∣B)P(A\mid B): the chance of an effect AA once the cause BB is known. Frequently we observe the effect and want the cause, P(B∣A)P(B\mid A). Example 11.5 makes this concrete. Urn I holds 22 green and 11 red ball; urn II holds 22 red and 33 yellow balls. We pick an urn at random and draw a ball; given that it is red, which urn did it come from? The forward probabilities P(red∣urn I)=13P(\text{red}\mid\text{urn I})=\tfrac13 and P(red∣urn II)=25P(\text{red}\mid\text{urn II})=\tfrac25 are immediate, but the reverse P(urn I∣red)P(\text{urn I}\mid\text{red}) is what we actually want. Rewrite it as P(urn I∣red)=P(red∩urn I)P(red)=P(red∣urn I)P(urn I)P(red)P(\text{urn I}\mid\text{red})=\dfrac{P(\text{red}\cap\text{urn I})}{P(\text{red})}=\dfrac{P(\text{red}\mid\text{urn I})P(\text{urn I})}{P(\text{red})} and expand P(red)P(\text{red}) with the law of total probability. The answer 511<12\tfrac{5}{11}<\tfrac12 says the red ball more likely came from urn II -- precisely because urn II is richer in red balls. Bayes' formula is the machine that performs this reversal in general.

2

Bayes' formula and where it comes from

Proposition 11.2 packages the reversal. Whenever P(A),P(B),P(Bc)>0P(A),P(B),P(B^c)>0, P(B∣A)=P(A∣B) P(B)P(A∣B) P(B)+P(A∣Bc) P(Bc).P(B\mid A)=\frac{P(A\mid B)\,P(B)}{P(A\mid B)\,P(B)+P(A\mid B^c)\,P(B^c)}. The derivation is two moves. First, the definition of conditional probability (Definition 11.1) used twice: P(B∣A)=P(A∩B)P(A)P(B\mid A)=\dfrac{P(A\cap B)}{P(A)} and P(A∩B)=P(A∣B)P(B)P(A\cap B)=P(A\mid B)P(B), so the numerator is P(A∣B)P(B)P(A\mid B)P(B). Second, the law of total probability (Proposition 11.1) splits the denominator across the partition {B,Bc}\{B,B^c\}: P(A)=P(A∣B)P(B)+P(A∣Bc)P(Bc)P(A)=P(A\mid B)P(B)+P(A\mid B^c)P(B^c). So the scary-looking denominator is just P(A)P(A) written in computable pieces. Everything required -- the two likelihoods P(A∣B),P(A∣Bc)P(A\mid B),P(A\mid B^c) and the prior P(B)P(B) -- is exactly the data a forward model already supplies.

3

Prior and posterior (Remark 11.2)

Remark 11.2 names the two ends of the computation. The prior probability P(B)P(B) is our belief before seeing evidence; the posterior probability P(B∣A)P(B\mid A) is that belief updated after observing AA. Bayes' formula is the engine that turns one into the other. Whether evidence raises or lowers belief depends on how diagnostic AA is: if AA is more likely under BB than under BcB^c (that is, P(A∣B)>P(A∣Bc)P(A\mid B)>P(A\mid B^c)) the posterior exceeds the prior; if it is less likely, the posterior drops; and in the knife-edge case P(A∣B)=P(A∣Bc)P(A\mid B)=P(A\mid B^c) the evidence is irrelevant and the posterior equals the prior -- which is exactly independence of AA and BB. The size of the shift, however, is still anchored by the prior, and that is the heart of the next idea.

4

The base-rate problem and the base-rate fallacy

The most famous use of Bayes is screening for a rare condition (Example 11.6). A test is described by its sensitivity P(+∣D)P(+\mid D) (true-positive rate), its false-positive rate P(+∣Dc)P(+\mid D^c) (so its specificity is 1−P(+∣Dc)1-P(+\mid D^c)), and the disease's prevalence P(D)P(D), which plays the role of the prior. Intuition whispers that a 96%96\%-sensitive test coming back positive means about a 96%96\% chance of disease. Bayes says otherwise. When the disease is rare, the enormous healthy population generates many false positives that swamp the few true positives, so the posterior P(D∣+)P(D\mid +) can sit far below 50%50\%. Reading the sensitivity as the answer -- ignoring the prior P(D)P(D) -- is the base-rate fallacy. The fix never changes: put the prevalence in the numerator and inside the total-probability denominator, and let the arithmetic correct the intuition.

Definition 11.1 -- Conditional Probability

For events A,BA,B with P(B)>0P(B)>0, the conditional probability of AA given BB is P(A∣B)=P(A∩B)P(B)P(A\mid B)=\dfrac{P(A\cap B)}{P(B)}. Clearing the denominator gives the multiplication rule P(A∩B)=P(A∣B) P(B)=P(B∣A) P(A)P(A\cap B)=P(A\mid B)\,P(B)=P(B\mid A)\,P(A).

Intuition. Conditioning on BB restricts attention to the outcomes inside BB and renormalizes so that their probabilities again sum to 11 -- hence the division by P(B)P(B). The multiplication rule is merely this definition with the denominator cleared, and it is the bridge that lets Bayes' formula trade P(A∩B)P(A\cap B) for the computable product P(A∣B)P(B)P(A\mid B)P(B).
Proposition 11.1 -- Law of Total Probability

If {B1,…,Bn}\{B_1,\dots,B_n\} is a partition of the sample space with every P(Bi)>0P(B_i)>0, then for any event AA, P(A)=∑i=1nP(A∩Bi)=∑i=1nP(A∣Bi) P(Bi)P(A)=\sum_{i=1}^n P(A\cap B_i)=\sum_{i=1}^n P(A\mid B_i)\,P(B_i). In particular, for a single event BB with 0<P(B)<10<P(B)<1, P(A)=P(A∣B)P(B)+P(A∣Bc)P(Bc)P(A)=P(A\mid B)P(B)+P(A\mid B^c)P(B^c).

Intuition. A partition slices the sample space into disjoint pieces, so the event AA is the disjoint union of its slivers A∩BiA\cap B_i and its probability is the sum of those pieces. Writing each piece as P(A∣Bi)P(Bi)P(A\mid B_i)P(B_i) computes P(A)P(A) from a forward model -- and this expression is exactly the denominator of Bayes' formula.
Proposition 11.2 -- Bayes' Formula

Let P(A),P(B),P(Bc)>0P(A),P(B),P(B^c)>0. Then P(B∣A)=P(A∣B) P(B)P(A∣B) P(B)+P(A∣Bc) P(Bc).P(B\mid A)=\frac{P(A\mid B)\,P(B)}{P(A\mid B)\,P(B)+P(A\mid B^c)\,P(B^c)}.

Intuition. Begin with P(B∣A)=P(A∩B)P(A)P(B\mid A)=\dfrac{P(A\cap B)}{P(A)} (Definition 11.1). Replace the numerator using the multiplication rule P(A∩B)=P(A∣B)P(B)P(A\cap B)=P(A\mid B)P(B), and expand the denominator P(A)P(A) with the law of total probability over {B,Bc}\{B,B^c\} (Proposition 11.1). The result reverses the conditioning, turning the forward probabilities P(A∣B),P(A∣Bc)P(A\mid B),P(A\mid B^c) and the prior P(B)P(B) into the reverse probability P(B∣A)P(B\mid A).
Remark 11.2 -- Prior and Posterior Probability

In Bayes' formula, P(B)P(B) is the prior probability (belief before the evidence) and P(B∣A)P(B\mid A) is the posterior probability (belief updated after observing AA).

Intuition. Bayes' formula is a belief-update rule: it maps the prior P(B)P(B) to the posterior P(B∣A)P(B\mid A) using the likelihoods P(A∣B)P(A\mid B) and P(A∣Bc)P(A\mid B^c). Evidence more probable under BB than under BcB^c pulls the posterior above the prior; equally probable evidence leaves it unchanged; less probable evidence pushes it down.

Worked examples

Example 1

Two urns (Example 11.5). Urn I contains 22 green and 11 red ball; urn II contains 22 red and 33 yellow balls. An urn is chosen at random (each with probability 12\tfrac12) and one ball is drawn from it. Given that the drawn ball is red, find the probability it came from urn I.

  1. 1

    Name the events. Let B={urn I}B=\{\text{urn I}\} (so Bc={urn II}B^c=\{\text{urn II}\}) and A={red}A=\{\text{red}\}. The priors are P(B)=P(Bc)=12P(B)=P(B^c)=\tfrac12.

  2. 2

    Read off the forward (likelihood) probabilities. Urn I has 11 red of 33 balls, so P(A∣B)=13P(A\mid B)=\tfrac13. Urn II has 22 red of 55 balls, so P(A∣Bc)=25P(A\mid B^c)=\tfrac25.

  3. 3

    Expand the denominator with the law of total probability (Proposition 11.1): P(A)=P(A∣B)P(B)+P(A∣Bc)P(Bc)=13⋅12+25⋅12=16+15=530+630=1130.P(A)=P(A\mid B)P(B)+P(A\mid B^c)P(B^c)=\tfrac13\cdot\tfrac12+\tfrac25\cdot\tfrac12=\tfrac16+\tfrac15=\tfrac{5}{30}+\tfrac{6}{30}=\tfrac{11}{30}.

  4. 4

    Apply Bayes' formula (Proposition 11.2): P(B∣A)=P(A∣B)P(B)P(A)=13⋅121130=161130=16⋅3011=511.P(B\mid A)=\dfrac{P(A\mid B)P(B)}{P(A)}=\dfrac{\tfrac13\cdot\tfrac12}{\tfrac{11}{30}}=\dfrac{\tfrac16}{\tfrac{11}{30}}=\tfrac16\cdot\tfrac{30}{11}=\tfrac{5}{11}.

  5. 5

    Interpret. Since 511≈0.455<12\tfrac{5}{11}\approx0.455<\tfrac12, the red ball more likely came from urn II, consistent with urn II holding more red balls. The prior 12\tfrac12 was revised down to the posterior 511\tfrac{5}{11}.

Answer. P(urn I∣red)=511≈0.455P(\text{urn I}\mid\text{red})=\dfrac{5}{11}\approx0.455. The posterior (0.4550.455) sits below the prior (0.50.5), so the red ball is more likely from urn II.
Example 2

Medical test -- the base-rate problem (Example 11.6). A test for a disease has sensitivity 96%96\% (positive on 96%96\% of people who have the disease) and a false-positive rate of 2%2\% (positive on 2%2\% of people who do not). The disease's prevalence is 0.5%0.5\% of the population. A randomly chosen person tests positive. What is the probability they actually have the disease?

Example 3

Spam filter. A filter has learned that the word free appears in 80%80\% of spam emails and in 10%10\% of legitimate (ham) emails. Suppose 40%40\% of all incoming email is spam. An email arrives containing the word free. Find the probability the email is spam, and identify the prior and the posterior.