Probability and Bayes: Uncertainty, Evidence, and Rational Updating

The mathematics of probability is precise, but its meaning is contested. Bayes' rule shows how conditional probabilities relate without turning uncertainty into truth.

Probability has an unusual philosophical structure. Its mathematics can be exact even when people disagree about what the number represents.

When someone says that an event has probability 0.7, they might be describing a long-run frequency, a physical tendency, the support supplied by evidence, or their own rational degree of confidence. The same formal rules can operate across these interpretations. The rules tell us how probabilities must fit together; they do not by themselves tell us what probabilities are.

Bayes’ rule works inside that formal structure. It relates an initial probability, new evidence, and an updated probability. It is indispensable for reasoning under uncertainty, but it does not guarantee that the starting assumptions, evidence, or model are correct.

This essay asks what a probability means and how Bayesian updating changes a judgment. Logic and Probability compares evidential support with entailment; Logic examines validity within inference itself.

One calculus, several meanings

Consider four claims:

A fair coin has a 50% probability of landing heads.
A radium atom has a certain probability of decaying within a year.
The evidence gives the defendant a 20% probability of guilt.
I am 80% confident that the train will arrive on time.

They all use probability language, but they do not obviously describe the same kind of property.

The first may be grounded in symmetry. The second appears to describe a physical process. The third concerns how evidence bears on a hypothesis. The fourth describes a person’s graded confidence.

The philosophy of probability asks what makes statements like these true or reasonable. The Stanford Encyclopedia of Philosophy groups the central possibilities around physical probability, evidential support, and degrees of confidence, while noting that their boundaries can overlap. Stanford Encyclopedia of Philosophy: Interpretations of Probability

The axioms constrain probability without interpreting it

A probability model begins with:

  • a sample space Ω, containing possible outcomes;
  • events represented as subsets of Ω;
  • a function P assigning numbers to those events.

The standard axioms require that probabilities are nonnegative, that P(Ω) = 1, and that mutually exclusive events add appropriately.

For an ordinary six-sided die:

Ω = {1, 2, 3, 4, 5, 6}
P(even) = P({2, 4, 6}) = 1/2

These rules establish coherence. They do not tell us why each face receives probability 1/6. That assignment might rely on physical symmetry, observed frequencies, a model of the die, or a state of information.

This separation is crucial:

Probability theory specifies relations among probability values. An interpretation explains what those values mean and how they should be assigned.

Chance in the world and uncertainty in knowledge

Some uncertainty appears to concern how the world behaves. Before a coin lands, the physical process may be modeled as chancy. Radioactive decay is commonly represented probabilistically even when the experimental conditions are carefully controlled.

Other uncertainty is plainly epistemic. A coin may already have landed under a cup. The result is fixed, but an observer who cannot see it may assign equal confidence to heads and tails.

This produces a persistent question:

Is probability a feature of the world,
a relation between evidence and propositions,
or a feature of an agent's information?

Frequency interpretations connect probability with proportions in repeated trials. Propensity accounts treat probability as a tendency or disposition of a setup to produce outcomes. Logical or evidential accounts connect probability with how strongly evidence supports a conclusion. Subjective or personalist accounts represent an agent’s coherent degree of belief.

No interpretation is automatically best for every use. A weather forecast, a quantum transition, a clinical trial, and a person’s confidence in a business decision may require different explanatory work even when they use the same calculus.

Conditional probability makes information explicit

Probabilities are often relative to conditions:

P(A | B)

reads as the probability of A given B. When P(B) > 0:

P(A | B) = P(A ∩ B) / P(B)

Suppose 40 of 100 people carry umbrellas, and 30 of those 40 arrive wet. Then:

P(wet | umbrella) = 30 / 40 = 0.75

This is not necessarily the same as:

P(umbrella | wet)

Reversing the condition changes the reference class. Much faulty probabilistic reasoning comes from treating these two quantities as interchangeable.

Bayes’ rule reverses a conditional relation

From the definition of conditional probability:

P(H ∩ E) = P(E | H)P(H)
P(H ∩ E) = P(H | E)P(E)

Therefore:

P(H | E) = P(E | H)P(H) / P(E)

Here:

  • H is a hypothesis;
  • E is observed evidence;
  • P(H) is the prior probability;
  • P(E | H) is the likelihood of the evidence under the hypothesis;
  • P(H | E) is the posterior probability.

The denominator can be expanded across competing hypotheses. If H and ¬H exhaust the possibilities:

P(E) = P(E | H)P(H) + P(E | ¬H)P(¬H)

Bayes’ rule is an identity. The controversy begins when it is used as a general account of learning: which priors are rational, which hypotheses belong in the model, and what should count as evidence?

Why base rates change the meaning of a positive test

Suppose a condition affects 1% of a population. A test has:

  • 90% sensitivity: P(positive | condition) = 0.90;
  • 5% false-positive rate: P(positive | no condition) = 0.05.

The probability of the condition after a positive result is:

P(condition | positive)
= (0.90 × 0.01) / [(0.90 × 0.01) + (0.05 × 0.99)]
= 0.009 / 0.0585
≈ 0.154

So the posterior probability is about 15.4%, not 90%.

A frequency representation makes this intuitive. Among 10,000 people:

  • about 100 have the condition, and 90 test positive;
  • about 9,900 do not, and roughly 495 test positive;
  • among 585 positive results, only 90 are true positives.

The test is informative: probability rises from 1% to about 15.4%. But sensitivity answers how often the test detects the condition when it is present. It does not directly answer how often the condition is present when the test is positive. Penn State STAT 414: Bayes’ Theorem

Bayesianism turns updating into a norm of belief

Bayesian epistemology represents confidence as a credence, a value between 0 and 1. It then asks how credences ought to fit together and how they ought to change when evidence arrives.

This approach separates two norms:

  1. Synchronic coherence: probabilities held at one time should satisfy the probability axioms.
  2. Diachronic updating: beliefs should change in a disciplined way as evidence changes.

Bayesian conditionalization is one proposed updating rule. If an agent becomes certain of evidence E, the new credence in H should often equal the old conditional credence P(H | E).

This gives a precise account of graded belief. It also explains why rationality need not require certainty. Two people can assign different probabilities while both responding coherently to evidence, especially when their prior information differs. Stanford Encyclopedia of Philosophy: Bayesian Epistemology

Updating cannot repair a bad model by itself

Bayesian reasoning is only as good as the space within which it updates.

Priors require justification

Some priors follow from measured base rates or well-tested models. Others express limited information, expert judgment, convention, or convenience. Labeling a number “prior” does not make it objective.

The hypothesis space may omit the truth

If a diagnosis system considers only three diseases while the patient has a fourth, all posterior probability will be redistributed among the wrong options.

Likelihoods depend on a model

P(E | H) assumes a relation between hypothesis and evidence. Measurement error, selection bias, dependence among observations, and changing environments can make the assumed likelihood unreliable.

Zero priors can block learning

If P(H) = 0, ordinary Bayesian updating leaves the posterior at zero regardless of the evidence. Absolute certainty assigned too early can make a model unable to recover.

Evidence selection is not neutral

Someone must decide what to measure, which observations are relevant, and how the data are encoded. Updating can be formally correct while the evidence pipeline remains biased or incomplete.

A posterior probability is not truth

A posterior expresses probability under a model and body of evidence. It is not a direct transformation of uncertainty into fact.

high posterior probability
≠ verified truth
≠ complete hypothesis space
≠ reliable measurement
≠ justified decision in every context

Decisions also depend on consequences. A 5% probability may justify action when the possible harm is catastrophic; a 95% probability may still be insufficient for an irreversible accusation.

Probability helps make uncertainty explicit. Bayes’ rule makes one important form of learning explicit. Their philosophical value lies in disciplined revision, not in a promise of certainty.

Sources

If this was useful, subscribe via RSS.

Content is open for citation with attribution; please link back to the source.