<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/"><channel><title>Bayes on Moonment</title><link>https://moonment.net/en/tags/bayes/</link><description>Moon's notes on concepts, real projects, and reasoning open to review.</description><generator>Hugo</generator><language>en-US</language><managingEditor>Moon</managingEditor><webMaster>Moon</webMaster><copyright>© 2026 Moonment</copyright><lastBuildDate>Tue, 29 Sep 2026 15:04:00 +0800</lastBuildDate><atom:link href="https://moonment.net/en/tags/bayes/index.xml" rel="self" type="application/rss+xml"/><item><title>Probability and Bayes: Uncertainty, Evidence, and Rational Updating</title><link>https://moonment.net/en/notes/probability-and-bayes/</link><pubDate>Sat, 12 Sep 2026 22:03:10 +0800</pubDate><dc:creator>Moon</dc:creator><guid>https://moonment.net/en/notes/probability-and-bayes/</guid><description>The mathematics of probability is precise, but its meaning is contested. Bayes' rule shows how conditional probabilities relate without turning uncertainty into truth.</description><content:encoded><![CDATA[<p>Probability has an unusual philosophical structure. Its mathematics can be exact even when people disagree about what the number represents.</p>
<p>When someone says that an event has probability 0.7, they might be describing a long-run frequency, a physical tendency, the support supplied by evidence, or their own rational degree of confidence. The same formal rules can operate across these interpretations. The rules tell us how probabilities must fit together; they do not by themselves tell us what probabilities are.</p>
<p>Bayes’ rule works inside that formal structure. It relates an initial probability, new evidence, and an updated probability. It is indispensable for reasoning under uncertainty, but it does not guarantee that the starting assumptions, evidence, or model are correct.</p>
<p>This essay asks what a probability means and how Bayesian updating changes a judgment. <a href="/en/notes/logic-and-probability/">Logic and Probability</a> compares evidential support with entailment; <a href="/en/notes/what-is-logic/">Logic</a> examines validity within inference itself.</p>
<h2 id="one-calculus-several-meanings">One calculus, several meanings</h2>
<p>Consider four claims:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">A fair coin has a 50% probability of landing heads.
</span></span><span class="line"><span class="cl">A radium atom has a certain probability of decaying within a year.
</span></span><span class="line"><span class="cl">The evidence gives the defendant a 20% probability of guilt.
</span></span><span class="line"><span class="cl">I am 80% confident that the train will arrive on time.
</span></span></code></pre></div><p>They all use probability language, but they do not obviously describe the same kind of property.</p>
<p>The first may be grounded in symmetry. The second appears to describe a physical process. The third concerns how evidence bears on a hypothesis. The fourth describes a person’s graded confidence.</p>
<p>The philosophy of probability asks what makes statements like these true or reasonable. The Stanford Encyclopedia of Philosophy groups the central possibilities around physical probability, evidential support, and degrees of confidence, while noting that their boundaries can overlap. <a href="https://plato.stanford.edu/entries/probability-interpret/">Stanford Encyclopedia of Philosophy: Interpretations of Probability</a></p>
<h2 id="the-axioms-constrain-probability-without-interpreting-it">The axioms constrain probability without interpreting it</h2>
<p>A probability model begins with:</p>
<ul>
<li>a sample space <code>Ω</code>, containing possible outcomes;</li>
<li>events represented as subsets of <code>Ω</code>;</li>
<li>a function <code>P</code> assigning numbers to those events.</li>
</ul>
<p>The standard axioms require that probabilities are nonnegative, that <code>P(Ω) = 1</code>, and that mutually exclusive events add appropriately.</p>
<p>For an ordinary six-sided die:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">Ω = {1, 2, 3, 4, 5, 6}
</span></span><span class="line"><span class="cl">P(even) = P({2, 4, 6}) = 1/2
</span></span></code></pre></div><p>These rules establish coherence. They do not tell us why each face receives probability <code>1/6</code>. That assignment might rely on physical symmetry, observed frequencies, a model of the die, or a state of information.</p>
<p>This separation is crucial:</p>
<blockquote>
<p><strong>Probability theory specifies relations among probability values. An interpretation explains what those values mean and how they should be assigned.</strong></p>
</blockquote>
<h2 id="chance-in-the-world-and-uncertainty-in-knowledge">Chance in the world and uncertainty in knowledge</h2>
<p>Some uncertainty appears to concern how the world behaves. Before a coin lands, the physical process may be modeled as chancy. Radioactive decay is commonly represented probabilistically even when the experimental conditions are carefully controlled.</p>
<p>Other uncertainty is plainly epistemic. A coin may already have landed under a cup. The result is fixed, but an observer who cannot see it may assign equal confidence to heads and tails.</p>
<p>This produces a persistent question:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">Is probability a feature of the world,
</span></span><span class="line"><span class="cl">a relation between evidence and propositions,
</span></span><span class="line"><span class="cl">or a feature of an agent&#39;s information?
</span></span></code></pre></div><p>Frequency interpretations connect probability with proportions in repeated trials. Propensity accounts treat probability as a tendency or disposition of a setup to produce outcomes. Logical or evidential accounts connect probability with how strongly evidence supports a conclusion. Subjective or personalist accounts represent an agent’s coherent degree of belief.</p>
<p>No interpretation is automatically best for every use. A weather forecast, a quantum transition, a clinical trial, and a person’s confidence in a business decision may require different explanatory work even when they use the same calculus.</p>
<h2 id="conditional-probability-makes-information-explicit">Conditional probability makes information explicit</h2>
<p>Probabilities are often relative to conditions:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">P(A | B)
</span></span></code></pre></div><p>reads as the probability of <code>A</code> given <code>B</code>. When <code>P(B) &gt; 0</code>:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">P(A | B) = P(A ∩ B) / P(B)
</span></span></code></pre></div><p>Suppose 40 of 100 people carry umbrellas, and 30 of those 40 arrive wet. Then:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">P(wet | umbrella) = 30 / 40 = 0.75
</span></span></code></pre></div><p>This is not necessarily the same as:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">P(umbrella | wet)
</span></span></code></pre></div><p>Reversing the condition changes the reference class. Much faulty probabilistic reasoning comes from treating these two quantities as interchangeable.</p>
<h2 id="bayes-rule-reverses-a-conditional-relation">Bayes’ rule reverses a conditional relation</h2>
<p>From the definition of conditional probability:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">P(H ∩ E) = P(E | H)P(H)
</span></span><span class="line"><span class="cl">P(H ∩ E) = P(H | E)P(E)
</span></span></code></pre></div><p>Therefore:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">P(H | E) = P(E | H)P(H) / P(E)
</span></span></code></pre></div><p>Here:</p>
<ul>
<li><code>H</code> is a hypothesis;</li>
<li><code>E</code> is observed evidence;</li>
<li><code>P(H)</code> is the prior probability;</li>
<li><code>P(E | H)</code> is the likelihood of the evidence under the hypothesis;</li>
<li><code>P(H | E)</code> is the posterior probability.</li>
</ul>
<p>The denominator can be expanded across competing hypotheses. If <code>H</code> and <code>¬H</code> exhaust the possibilities:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">P(E) = P(E | H)P(H) + P(E | ¬H)P(¬H)
</span></span></code></pre></div><p>Bayes’ rule is an identity. The controversy begins when it is used as a general account of learning: which priors are rational, which hypotheses belong in the model, and what should count as evidence?</p>
<h2 id="why-base-rates-change-the-meaning-of-a-positive-test">Why base rates change the meaning of a positive test</h2>
<p>Suppose a condition affects 1% of a population. A test has:</p>
<ul>
<li>90% sensitivity: <code>P(positive | condition) = 0.90</code>;</li>
<li>5% false-positive rate: <code>P(positive | no condition) = 0.05</code>.</li>
</ul>
<p>The probability of the condition after a positive result is:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">P(condition | positive)
</span></span><span class="line"><span class="cl">= (0.90 × 0.01) / [(0.90 × 0.01) + (0.05 × 0.99)]
</span></span><span class="line"><span class="cl">= 0.009 / 0.0585
</span></span><span class="line"><span class="cl">≈ 0.154
</span></span></code></pre></div><p>So the posterior probability is about 15.4%, not 90%.</p>
<p>A frequency representation makes this intuitive. Among 10,000 people:</p>
<ul>
<li>about 100 have the condition, and 90 test positive;</li>
<li>about 9,900 do not, and roughly 495 test positive;</li>
<li>among 585 positive results, only 90 are true positives.</li>
</ul>
<p>The test is informative: probability rises from 1% to about 15.4%. But sensitivity answers how often the test detects the condition when it is present. It does not directly answer how often the condition is present when the test is positive. <a href="https://online.stat.psu.edu/stat414/Lesson06">Penn State STAT 414: Bayes’ Theorem</a></p>
<h2 id="bayesianism-turns-updating-into-a-norm-of-belief">Bayesianism turns updating into a norm of belief</h2>
<p>Bayesian epistemology represents confidence as a <strong>credence</strong>, a value between 0 and 1. It then asks how credences ought to fit together and how they ought to change when evidence arrives.</p>
<p>This approach separates two norms:</p>
<ol>
<li><strong>Synchronic coherence:</strong> probabilities held at one time should satisfy the probability axioms.</li>
<li><strong>Diachronic updating:</strong> beliefs should change in a disciplined way as evidence changes.</li>
</ol>
<p>Bayesian conditionalization is one proposed updating rule. If an agent becomes certain of evidence <code>E</code>, the new credence in <code>H</code> should often equal the old conditional credence <code>P(H | E)</code>.</p>
<p>This gives a precise account of graded belief. It also explains why rationality need not require certainty. Two people can assign different probabilities while both responding coherently to evidence, especially when their prior information differs. <a href="https://plato.stanford.edu/entries/epistemology-bayesian/">Stanford Encyclopedia of Philosophy: Bayesian Epistemology</a></p>
<h2 id="updating-cannot-repair-a-bad-model-by-itself">Updating cannot repair a bad model by itself</h2>
<p>Bayesian reasoning is only as good as the space within which it updates.</p>
<h3 id="priors-require-justification">Priors require justification</h3>
<p>Some priors follow from measured base rates or well-tested models. Others express limited information, expert judgment, convention, or convenience. Labeling a number “prior” does not make it objective.</p>
<h3 id="the-hypothesis-space-may-omit-the-truth">The hypothesis space may omit the truth</h3>
<p>If a diagnosis system considers only three diseases while the patient has a fourth, all posterior probability will be redistributed among the wrong options.</p>
<h3 id="likelihoods-depend-on-a-model">Likelihoods depend on a model</h3>
<p><code>P(E | H)</code> assumes a relation between hypothesis and evidence. Measurement error, selection bias, dependence among observations, and changing environments can make the assumed likelihood unreliable.</p>
<h3 id="zero-priors-can-block-learning">Zero priors can block learning</h3>
<p>If <code>P(H) = 0</code>, ordinary Bayesian updating leaves the posterior at zero regardless of the evidence. Absolute certainty assigned too early can make a model unable to recover.</p>
<h3 id="evidence-selection-is-not-neutral">Evidence selection is not neutral</h3>
<p>Someone must decide what to measure, which observations are relevant, and how the data are encoded. Updating can be formally correct while the evidence pipeline remains biased or incomplete.</p>
<h2 id="a-posterior-probability-is-not-truth">A posterior probability is not truth</h2>
<p>A posterior expresses probability under a model and body of evidence. It is not a direct transformation of uncertainty into fact.</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">high posterior probability
</span></span><span class="line"><span class="cl">≠ verified truth
</span></span><span class="line"><span class="cl">≠ complete hypothesis space
</span></span><span class="line"><span class="cl">≠ reliable measurement
</span></span><span class="line"><span class="cl">≠ justified decision in every context
</span></span></code></pre></div><p>Decisions also depend on consequences. A 5% probability may justify action when the possible harm is catastrophic; a 95% probability may still be insufficient for an irreversible accusation.</p>
<p>Probability helps make uncertainty explicit. Bayes’ rule makes one important form of learning explicit. Their philosophical value lies in disciplined revision, not in a promise of certainty.</p>
<h2 id="sources">Sources</h2>
<ul>
<li><a href="https://plato.stanford.edu/entries/probability-interpret/">Stanford Encyclopedia of Philosophy: Interpretations of Probability</a></li>
<li><a href="https://plato.stanford.edu/entries/epistemology-bayesian/">Stanford Encyclopedia of Philosophy: Bayesian Epistemology</a></li>
<li><a href="https://online.stat.psu.edu/stat414/Lesson06">Penn State STAT 414: Bayes’ Theorem</a></li>
</ul>
]]></content:encoded></item></channel></rss>