Unit content
Bayes' theorem
Conditional probability can be read in two directions. We may know how likely some evidence $B$ is under a hypothesis $A$, while what we actually want is the probability of $A$ after observing $B$.
Bayes' theorem reverses that conditioning:
$$P(A\mid B)=\frac{P(B\mid A)P(A)}{P(B)},\qquad P(B)>0.$$
Prior, likelihood and posterior
The quantity $P(A)$ is the prior probability of $A$. The conditional probability $P(B\mid A)$ describes how compatible the evidence is with $A$. After observing $B$, the updated probability
$$P(A\mid B)$$
is the posterior probability.
Expanding the denominator
If $A_1,\ldots,A_n$ form a partition of the sample space, then
$$P(B)=\sum_i P(B\mid A_i)P(A_i).$$
This makes Bayes' theorem practical when the evidence could have arisen under several mutually exclusive cases.
Diagnostic-test example
Suppose a disease has prevalence $1%$, a test detects it $99%$ of the time when present, and gives a false positive $5%$ of the time when absent. For disease event $D$ and positive result $+$,
$$P(D\mid +) =\frac{0.99\cdot0.01}{0.99\cdot0.01+0.05\cdot0.99} \approx0.167.$$
A highly sensitive test can therefore still have a modest posterior probability when the condition is rare.
Bayes' theorem updates uncertainty using evidence; it does not by itself establish a causal relationship.