<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/"><channel><title>Probability on Moonment</title><link>https://moonment.net/en/tags/probability/</link><description>Moon's notes on concepts, real projects, and reasoning open to review.</description><generator>Hugo</generator><language>en-US</language><managingEditor>Moon</managingEditor><webMaster>Moon</webMaster><copyright>© 2026 Moonment</copyright><lastBuildDate>Tue, 29 Sep 2026 15:04:00 +0800</lastBuildDate><atom:link href="https://moonment.net/en/tags/probability/index.xml" rel="self" type="application/rss+xml"/><item><title>Logic and Probability: Deduction, Uncertainty, and Evidence</title><link>https://moonment.net/en/notes/logic-and-probability/</link><pubDate>Sun, 27 Sep 2026 23:27:00 +0800</pubDate><dc:creator>Moon</dc:creator><guid>https://moonment.net/en/notes/logic-and-probability/</guid><description>Logic constrains what follows from premises; probability represents uncertainty and evidential support. This essay separates truth, validity, credence, conditional probability, Bayes, causation, and AI generation.</description><content:encoded><![CDATA[<p>Logic and probability both discipline inference, but they do not ask the same question.</p>
<blockquote>
<p><strong>Logic asks what follows from what. Probability asks how strongly the available information supports competing possibilities.</strong></p>
</blockquote>
<p>That distinction matters whenever evidence is incomplete. A conclusion can be logically valid but based on false premises. A hypothesis can be strongly supported without being entailed. A probability can equal one inside a model without expressing a logical truth.</p>
<p>Logic provides structure. Probability represents uncertainty within a structure. Neither can replace the other.</p>
<p>This essay focuses on their interface: why entailment is not conditional probability, and how deductive consequence relates to graded evidential support. <a href="/en/notes/what-is-logic/">Logic</a> treats consequence in its own right; <a href="/en/notes/probability-and-bayes/">Probability and Bayes</a> examines interpretations of probability and belief revision.</p>
<h2 id="the-scope-of-logic">The scope of “logic”</h2>
<p>Logic includes many systems: classical and non-classical logics, modal logic, temporal logic, inductive logic, and accounts of defeasible reasoning. The clearest starting point for comparison is classical deductive logic.</p>
<p>Classical logic studies propositions, truth values, and consequence. An argument is valid when there is no interpretation in which all its premises are true and its conclusion is false.<a href="https://plato.stanford.edu/entries/logic-classical/">Stanford Encyclopedia of Philosophy: Classical Logic</a></p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">All humans are mortal.
</span></span><span class="line"><span class="cl">Socrates is human.
</span></span><span class="line"><span class="cl">Therefore Socrates is mortal.
</span></span></code></pre></div><p>If both premises are true, the conclusion cannot be false. The relation can be written:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">D ⊨ C
</span></span></code></pre></div><p>This says that every interpretation satisfying premises <code>D</code> also satisfies conclusion <code>C</code>.</p>
<h2 id="validity-truth-and-soundness">Validity, truth, and soundness</h2>
<p>Validity concerns the form of an inference. It does not verify the premises.</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">All fish can fly.
</span></span><span class="line"><span class="cl">Carp are fish.
</span></span><span class="line"><span class="cl">Therefore carp can fly.
</span></span></code></pre></div><p>The form is valid. The first premise is false. A sound argument therefore requires both:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">valid inference
</span></span><span class="line"><span class="cl">+ true premises
</span></span></code></pre></div><p>This produces three separate questions:</p>
<ol>
<li>Are the concepts and propositions clear?</li>
<li>Are the premises true or adequately supported?</li>
<li>Does the conclusion follow from them?</li>
</ol>
<p>Probability often enters the second question. Evidence may support a premise to some degree even when it cannot establish it deductively.</p>
<h2 id="what-probability-represents">What probability represents</h2>
<p>Probability assigns values between zero and one to events or propositions, but the meaning of those values depends on interpretation.</p>
<p>Probability may represent:</p>
<ul>
<li>long-run frequency across repeated trials;</li>
<li>an objective chance or propensity in a physical system;</li>
<li>evidential support for a proposition;</li>
<li>a rational or personal degree of belief;</li>
<li>the output distribution of a statistical model.</li>
</ul>
<p>These interpretations share mathematical rules without making the same philosophical claim about what probability is.<a href="https://plato.stanford.edu/entries/probability-interpret/">Stanford Encyclopedia of Philosophy: Interpretations of Probability</a></p>
<p>“There is a 70% probability of rain tomorrow” may summarize a calibrated forecast over comparable cases, a model distribution, or a degree of belief given current evidence. It does not say that rain is logically required.</p>
<h2 id="two-different-relations">Two different relations</h2>
<table>
  <thead>
      <tr>
          <th>Question</th>
          <th>Logic</th>
          <th>Probability</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Central concern</td>
          <td>Does the conclusion follow from the premises?</td>
          <td>How much support does the evidence give a possibility?</td>
      </tr>
      <tr>
          <td>Typical expression</td>
          <td>If A, then B</td>
          <td><code>P(B | A) = 0.7</code></td>
      </tr>
      <tr>
          <td>Strength</td>
          <td>necessary, possible, impossible</td>
          <td>a degree from 0 to 1</td>
      </tr>
      <tr>
          <td>Main failures</td>
          <td>contradiction, invalid inference, equivocation</td>
          <td>bad conditioning, ignored base rates, misspecified models</td>
      </tr>
      <tr>
          <td>Response to new information</td>
          <td>add, remove, or revise premises</td>
          <td>update a probability distribution</td>
      </tr>
  </tbody>
</table>
<p>Logical consequence is categorical relative to the premises:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">A ⊨ B
</span></span></code></pre></div><p>Conditional probability is graded:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">P(B | A) = 0.9
</span></span></code></pre></div><p>The second expression still allows cases in which A is true and B is false. A high conditional probability is not an entailment.</p>
<h2 id="truth-is-not-a-probability-value">Truth is not a probability value</h2>
<p>In classical logic, a proposition under an interpretation is true or false. Probability describes uncertainty about events or propositions; it does not turn truth into a percentage.</p>
<p>Before tomorrow arrives, a forecast may assign a 70% probability to rain. After time, place, and the criterion for rain are fixed, the proposition “it rained” is either true or false. The earlier probability described an uncertain epistemic or predictive state.</p>
<p>It helps to distinguish:</p>
<table>
  <thead>
      <tr>
          <th>Level</th>
          <th>Question</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>truth</td>
          <td>Is the proposition actually the case?</td>
      </tr>
      <tr>
          <td>evidential support</td>
          <td>How strongly does the available evidence support it?</td>
      </tr>
      <tr>
          <td>credence</td>
          <td>How strongly does an agent believe it?</td>
      </tr>
  </tbody>
</table>
<p>Evidence and credence can be represented probabilistically. Neither is identical to truth.</p>
<h2 id="probability-one-is-not-always-logical-necessity">Probability one is not always logical necessity</h2>
<p>If <code>D</code> logically entails <code>C</code>, and <code>P(D) &gt; 0</code>, a probability model that respects the logical relation must satisfy:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">D ⊨ C
</span></span><span class="line"><span class="cl">→ P(C | D) = 1
</span></span></code></pre></div><p>The converse does not generally hold:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">P(C | D) = 1
</span></span><span class="line"><span class="cl">⇏ D ⊨ C
</span></span></code></pre></div><p>Probability one means that the model assigns all relevant probability mass to the event. Logical necessity means that no interpretation satisfying the premises makes the proposition false.</p>
<p>Continuous distributions make the difference vivid. A single exact point can have probability zero while remaining a possible value. Probability zero therefore need not mean contradiction, just as probability one need not mean logical truth.</p>
<h2 id="invalid-deduction-can-still-contain-evidence">Invalid deduction can still contain evidence</h2>
<p>Consider:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">If it rains, the ground becomes wet.
</span></span><span class="line"><span class="cl">The ground is wet.
</span></span><span class="line"><span class="cl">Therefore it rained.
</span></span></code></pre></div><p>As a deductive argument, this affirms the consequent and is invalid. Sprinklers, cleaning, or a leak could also wet the ground.</p>
<p>Yet wet ground may raise the probability of rain when:</p>
<ul>
<li>rain nearly always wets the ground;</li>
<li>other causes of wet ground are uncommon;</li>
<li>rain itself is not extremely rare.</li>
</ul>
<p>The observation can support the hypothesis without proving it:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">not deductively entailed
</span></span><span class="line"><span class="cl">but probabilistically confirmed
</span></span></code></pre></div><p>Inductive logic studies relations of this kind: premises may make a conclusion more credible without guaranteeing it.<a href="https://plato.stanford.edu/entries/logic-inductive/">Stanford Encyclopedia of Philosophy: Inductive Logic</a></p>
<h2 id="probability-depends-on-logical-structure">Probability depends on logical structure</h2>
<p>Probabilities cannot be assigned coherently until the events or propositions are specified.</p>
<p>One must know:</p>
<ul>
<li>which events exclude one another;</li>
<li>which can occur together;</li>
<li>whether one event includes another;</li>
<li>what the condition in a conditional probability means;</li>
<li>what counts as the negation of an event;</li>
<li>whether the listed possibilities are exhaustive.</li>
</ul>
<p>Suppose:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">A = a user clicked an advertisement
</span></span><span class="line"><span class="cl">B = a user completed a purchase attributed to that click
</span></span></code></pre></div><p>If the operational definition makes B a subset of A, then:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">B → A
</span></span><span class="line"><span class="cl">P(B) ≤ P(A)
</span></span></code></pre></div><p>A report showing more attributed buyers than recorded clickers signals a definition, attribution, collection, or data-integration problem. A more sophisticated probability formula will not repair an incoherent event structure.</p>
<h2 id="bayes-connects-evidence-and-belief-revision">Bayes connects evidence and belief revision</h2>
<p>Bayes&rsquo; theorem is:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">P(H | E) = P(E | H) × P(H) / P(E)
</span></span></code></pre></div><p>Here:</p>
<ul>
<li><code>H</code> is a hypothesis;</li>
<li><code>E</code> is evidence;</li>
<li><code>P(H)</code> is the prior probability;</li>
<li><code>P(E | H)</code> is the likelihood of the evidence if the hypothesis is true;</li>
<li><code>P(H | E)</code> is the posterior probability after observing the evidence.</li>
</ul>
<p>Bayesian reasoning does not assert:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">E occurred
</span></span><span class="line"><span class="cl">→ H must be true
</span></span></code></pre></div><p>It compares how expected the evidence would be under rival hypotheses, then reallocates confidence. Logical relations define hypotheses, evidence, exclusions, and implications. Probability quantifies the resulting uncertainty. Bayesian epistemology develops this into a normative account of rational belief revision.<a href="https://plato.stanford.edu/entries/epistemology-bayesian/">Stanford Encyclopedia of Philosophy: Bayesian Epistemology</a></p>
<p>Bayes also exposes a common error: confusing <code>P(E | H)</code> with <code>P(H | E)</code>. A test may be highly likely to return positive when a condition is present while the probability of the condition given a positive result remains much lower, especially when the condition is rare.</p>
<h2 id="probability-is-not-causation">Probability is not causation</h2>
<p>Logic, probability, and causation answer different questions:</p>
<table>
  <thead>
      <tr>
          <th>Relation</th>
          <th>Question</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>logical</td>
          <td>What must be accepted if the premises are accepted?</td>
      </tr>
      <tr>
          <td>probabilistic</td>
          <td>How does conditioning on information change uncertainty?</td>
      </tr>
      <tr>
          <td>causal</td>
          <td>What would change under an intervention, and through what process?</td>
      </tr>
  </tbody>
</table>
<p>A strong association may arise from reverse causation, a common cause, selection, measurement, or random variation. Causal analysis adds temporal order, counterfactual comparisons, interventions, mechanisms, and assumptions that identify an effect. The fuller account is developed in <a href="/en/notes/causality-causes-and-reasons/">What Causation Means</a>.</p>
<h2 id="probability-does-not-choose-an-action">Probability does not choose an action</h2>
<p>A well-calibrated probability still leaves practical questions unresolved:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">logic: is the reasoning coherent?
</span></span><span class="line"><span class="cl">probability: how likely are the outcomes?
</span></span><span class="line"><span class="cl">value: how good or bad are the outcomes?
</span></span><span class="line"><span class="cl">risk: which losses are tolerable?
</span></span><span class="line"><span class="cl">authority: who may make the choice?
</span></span><span class="line"><span class="cl">decision: which action is selected?
</span></span></code></pre></div><p>The option with the highest probability of success may have a trivial benefit, an unacceptable downside, or costs imposed on people who did not authorize the decision. Probability supplies inputs to decision-making; it does not settle values and responsibility.</p>
<h2 id="logic-and-probability-in-ai-systems">Logic and probability in AI systems</h2>
<p>A language model assigns probabilities to possible next tokens given context, then a decoding procedure selects outputs:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">context
</span></span><span class="line"><span class="cl">→ probability distribution over next tokens
</span></span><span class="line"><span class="cl">→ token selection
</span></span><span class="line"><span class="cl">→ generated text
</span></span></code></pre></div><p>High generation probability does not establish that a sentence is true, logically entailed, responsive to the user&rsquo;s actual aim, or authorized for action.</p>
<p>An AI system therefore needs more than probabilistic generation. Depending on the task, it may need:</p>
<ul>
<li>factual retrieval and source checks;</li>
<li>consistency and schema validation;</li>
<li>explicit rules and permission checks;</li>
<li>calculations or formal proofs;</li>
<li>execution results and external feedback.</li>
</ul>
<p>A fluent answer may be probable but contradictory. A valid derivation may be built on false retrieved facts. A calibrated prediction may still identify no useful intervention. These are different failure modes and require different checks.</p>
<h2 id="an-audit-for-uncertain-inference">An audit for uncertain inference</h2>
<p>When reading or constructing an argument under uncertainty, ask:</p>
<ol>
<li>What exactly are the propositions or events?</li>
<li>Which statements are premises, observations, assumptions, or definitions?</li>
<li>Is the conclusion entailed or only supported to a degree?</li>
<li>What evidence supports the premises?</li>
<li>What interpretation does the probability number have?</li>
<li>Is the conditioning information stated correctly?</li>
<li>Have base rates and rival hypotheses been considered?</li>
<li>Has an association or prediction been mistaken for a cause?</li>
<li>Which values, risks, and permissions remain outside the probability model?</li>
<li>What new evidence would change the conclusion?</li>
</ol>
<h2 id="conclusion">Conclusion</h2>
<p>Logic and probability impose different kinds of discipline on reasoning.</p>
<blockquote>
<p><strong>Logic specifies constraints among propositions and identifies what follows from accepted premises. Probability represents uncertainty about events or propositions and constrains how confidence should respond to evidence.</strong></p>
</blockquote>
<p>Their connection can be summarized as:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">logic defines the structure
</span></span><span class="line"><span class="cl">→ probability represents uncertainty within it
</span></span><span class="line"><span class="cl">→ evidence updates probabilities
</span></span><span class="line"><span class="cl">→ causal inquiry asks what changes what
</span></span><span class="line"><span class="cl">→ values and risks enter decisions
</span></span><span class="line"><span class="cl">→ action produces new evidence
</span></span></code></pre></div><p>Logic cannot replace probability when evidence is incomplete. Probability cannot replace logic when definitions conflict, possibilities are omitted, or an inference is invalid. Sound reasoning requires both the structure of consequence and the discipline of uncertainty.</p>
<h2 id="references">References</h2>
<ul>
<li><a href="https://plato.stanford.edu/entries/logic-classical/">Stanford Encyclopedia of Philosophy: Classical Logic</a></li>
<li><a href="https://plato.stanford.edu/entries/logical-consequence/">Stanford Encyclopedia of Philosophy: Logical Consequence</a></li>
<li><a href="https://plato.stanford.edu/entries/probability-interpret/">Stanford Encyclopedia of Philosophy: Interpretations of Probability</a></li>
<li><a href="https://plato.stanford.edu/entries/logic-inductive/">Stanford Encyclopedia of Philosophy: Inductive Logic</a></li>
<li><a href="https://plato.stanford.edu/entries/epistemology-bayesian/">Stanford Encyclopedia of Philosophy: Bayesian Epistemology</a></li>
</ul>
]]></content:encoded></item><item><title>Expected Value: Probability, Risk, and Decision</title><link>https://moonment.net/en/notes/what-is-expected-value/</link><pubDate>Sun, 27 Sep 2026 23:00:28 +0800</pubDate><dc:creator>Moon</dc:creator><guid>https://moonment.net/en/notes/what-is-expected-value/</guid><description>Expected value is a probability-weighted mean, not a prediction of the next outcome. This essay explains its mathematics, uses, limits, relation to risk, and role in decision theory.</description><content:encoded><![CDATA[<p>Expected value is often described as “what you can expect.” That phrase is convenient and dangerous. The expected value of a gamble may be an outcome that can never occur. It need not be the most likely outcome, and it does not promise what will happen next.</p>
<p>Expected value is a mathematical property of a probability distribution:</p>
<blockquote>
<p><strong>It is the probability-weighted mean of the values taken by a random variable.</strong></p>
</blockquote>
<p>Its importance comes from combining consequences and probabilities in one quantity. Its limitation is exactly the same: a single mean cannot preserve the full shape of a distribution or decide what risks a particular agent should accept.</p>
<h2 id="random-variables-and-distributions">Random variables and distributions</h2>
<p>A random variable assigns numerical values to outcomes. For a fair coin, define:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">heads → X = 1
</span></span><span class="line"><span class="cl">tails → X = 0
</span></span></code></pre></div><p>The corresponding distribution is:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">P(X = 1) = 0.5
</span></span><span class="line"><span class="cl">P(X = 0) = 0.5
</span></span></code></pre></div><p>For a discrete random variable, expected value is defined as:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">E[X] = Σ xᵢ P(X = xᵢ)
</span></span></code></pre></div><p>For the coin:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">E[X] = 1 × 0.5 + 0 × 0.5 = 0.5
</span></span></code></pre></div><p>No single toss produces half a head. The expectation belongs to the distribution, not to an individual trial. OpenStax therefore describes expected value as the mean of a discrete random variable and, under repeated trials, as its long-run average.<a href="https://openstax.org/books/statistics/pages/4-2-mean-or-expected-value-and-standard-deviation">OpenStax: Mean or Expected Value and Standard Deviation</a></p>
<h2 id="expected-value-is-not-the-most-likely-outcome">Expected value is not the most likely outcome</h2>
<p>Consider a lottery:</p>
<table>
  <thead>
      <tr>
          <th>Outcome</th>
          <th style="text-align: right">Probability</th>
          <th style="text-align: right">Payoff</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>win</td>
          <td style="text-align: right">10%</td>
          <td style="text-align: right">$100</td>
      </tr>
      <tr>
          <td>lose</td>
          <td style="text-align: right">90%</td>
          <td style="text-align: right">$0</td>
      </tr>
  </tbody>
</table>
<p>Its expected payoff is:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">E[X] = 0.1 × 100 + 0.9 × 0 = 10
</span></span></code></pre></div><p>Yet $10 is not a possible payoff, and $0 is the most likely result.</p>
<p>Several summaries answer different questions:</p>
<table>
  <thead>
      <tr>
          <th>Quantity</th>
          <th>Question</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>expected value</td>
          <td>Where is the probability-weighted mean?</td>
      </tr>
      <tr>
          <td>mode</td>
          <td>Which outcome is most likely?</td>
      </tr>
      <tr>
          <td>median</td>
          <td>Which value divides the probability mass in half?</td>
      </tr>
      <tr>
          <td>variance</td>
          <td>How widely do outcomes spread around the mean?</td>
      </tr>
      <tr>
          <td>quantile</td>
          <td>What threshold contains a stated share of outcomes?</td>
      </tr>
      <tr>
          <td>worst case</td>
          <td>How severe can the loss become?</td>
      </tr>
  </tbody>
</table>
<p>Two distributions can have the same expected value while assigning radically different probabilities to gains, losses, and extreme outcomes.</p>
<h2 id="when-does-expectation-become-a-long-run-average">When does expectation become a long-run average?</h2>
<p>Expected value is defined from a distribution. Its interpretation as an observed long-run average requires additional conditions.</p>
<p>The intuitive story assumes that:</p>
<ul>
<li>comparable trials can be repeated;</li>
<li>the generating process remains stable;</li>
<li>observations have suitable independence or regularity;</li>
<li>the expectation exists and is finite;</li>
<li>the agent can remain in the process long enough for averaging to matter.</li>
</ul>
<p>Under appropriate conditions, averages across many trials can approach the expected value. That does not imply that a single observation should be close to it.</p>
<p>Many important choices are not indefinitely repeatable. A medical intervention, an irreversible project, or a decision that can exhaust all available capital may give one agent only one relevant draw. Expected value can still describe the modeled distribution, but “it works on average” is not a complete personal decision rule.</p>
<h2 id="why-expected-value-is-useful">Why expected value is useful</h2>
<p>Expected value compresses a distribution into a quantity that can be compared and combined. It is especially useful when:</p>
<ul>
<li>decisions repeat many times;</li>
<li>losses can be pooled or diversified;</li>
<li>outcomes have a common numerical scale;</li>
<li>probabilities are reasonably stable;</li>
<li>no single adverse outcome destroys the decision maker.</li>
</ul>
<p>Expectation is linear:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">E[aX + bY] = aE[X] + bE[Y]
</span></span></code></pre></div><p>This property does not require <code>X</code> and <code>Y</code> to be independent. It allows the expected value of a total to be assembled from the expectations of its components, which makes expectation central in statistics, finance, insurance, operations research, and machine learning.</p>
<h2 id="equal-means-can-hide-unequal-risks">Equal means can hide unequal risks</h2>
<p>Compare two choices:</p>
<table>
  <thead>
      <tr>
          <th>Choice</th>
          <th>Outcome</th>
          <th style="text-align: right">Expected value</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>A</td>
          <td>receive $10 for certain</td>
          <td style="text-align: right">$10</td>
      </tr>
      <tr>
          <td>B</td>
          <td>50% receive $100; 50% lose $80</td>
          <td style="text-align: right">$10</td>
      </tr>
  </tbody>
</table>
<p>Their expected values are identical. Their distributions are not.</p>
<p>A decision maker still needs to inspect:</p>
<ol>
<li><strong>Dispersion:</strong> How far can outcomes depart from the mean?</li>
<li><strong>Tail risk:</strong> Can a low-probability loss be catastrophic?</li>
<li><strong>Ruin:</strong> Can one failure remove the ability to continue?</li>
<li><strong>Timing:</strong> When do costs and benefits occur?</li>
<li><strong>Reversibility:</strong> Can the choice be undone or repeated?</li>
<li><strong>Dependence:</strong> Do losses arrive together rather than independently?</li>
<li><strong>Model error:</strong> How reliable are the estimated probabilities and values?</li>
</ol>
<p>Expected value is one feature of a distribution. Treating it as the distribution itself discards the information most relevant to many high-stakes choices.</p>
<h2 id="from-expected-value-to-expected-utility">From expected value to expected utility</h2>
<p>A dollar does not have the same practical significance in every state or for every person. Losing $10,000 may be tolerable for one agent and ruinous for another. The numerical payoff and its value to the decision maker must therefore be distinguished.</p>
<p>Expected utility applies a utility function to outcomes before averaging:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">E[U(X)] = Σ P(X = xᵢ)U(xᵢ)
</span></span></code></pre></div><p>In standard normative decision theory, an option is evaluated by combining beliefs about possible outcomes with the agent&rsquo;s valuation of those outcomes. Under specified consistency conditions, preferences can be represented as maximizing expected utility.<a href="https://plato.stanford.edu/entries/decision-theory/">Stanford Encyclopedia of Philosophy: Decision Theory</a></p>
<p>This is a normative representation of rational choice under uncertainty, not a complete psychological description of how people actually decide. Kahneman and Tversky developed prospect theory partly as a descriptive challenge to expected utility accounts. Their experiments emphasized reference dependence, the special weight of certainty, and different patterns of risk attitude for gains and losses.<a href="https://www.ucl.ac.uk/anaesthesia/sites/anaesthesia/files/kahneman-tversky.pdf">Kahneman and Tversky: Prospect Theory</a></p>
<h2 id="what-do-the-probabilities-mean">What do the probabilities mean?</h2>
<p>Every expected value inherits the interpretation and quality of its probabilities. A probability may represent:</p>
<ul>
<li>a long-run frequency;</li>
<li>an objective physical chance or propensity;</li>
<li>evidential support for a proposition;</li>
<li>an agent&rsquo;s degree of belief;</li>
<li>a predictive distribution estimated by a model.</li>
</ul>
<p>These interpretations are related but not interchangeable. The philosophy of probability distinguishes physical, evidential, and subjective readings, each of which changes what an expected value claim means.<a href="https://plato.stanford.edu/entries/probability-interpret/">Stanford Encyclopedia of Philosophy: Interpretations of Probability</a></p>
<p>A calculation may be arithmetically exact while its inputs are poor. Outcomes may have been omitted, values may be measured on the wrong scale, or probabilities may come from stale data and an incorrect model.</p>
<blockquote>
<p><strong>An expected value is no more reliable than its outcome definitions, value assignments, and probability estimates.</strong></p>
</blockquote>
<h2 id="expected-value-does-not-choose-by-itself">Expected value does not choose by itself</h2>
<p>To use expected value in a real decision:</p>
<ol>
<li>Define the available actions.</li>
<li>List materially different outcomes for each action.</li>
<li>Check for omitted indirect effects.</li>
<li>State where each probability comes from.</li>
<li>Decide what numerical value is being measured.</li>
<li>Compute the expectation.</li>
<li>Inspect dispersion, quantiles, dependence, and tails.</li>
<li>Test whether the agent can survive the downside.</li>
<li>Vary uncertain inputs and see whether the ranking changes.</li>
<li>Update the model when new evidence arrives.</li>
</ol>
<p>The expected value is an input to judgment. It cannot determine whether the modeled objective is morally acceptable, whether a loss is survivable, or whether the decision maker has authority to expose others to the risk.</p>
<h2 id="a-final-definition">A final definition</h2>
<blockquote>
<p><strong>Expected value is the probability-weighted mean of a random variable, used to summarize the center of its probability distribution.</strong></p>
</blockquote>
<p>It is not:</p>
<ul>
<li>a promise about the next observation;</li>
<li>the most likely outcome;</li>
<li>necessarily a value that can occur;</li>
<li>a full description of risk;</li>
<li>an automatic decision.</li>
</ul>
<p>Expected value makes uncertain consequences comparable. Good judgment begins after that calculation, by restoring the information that the average leaves out.</p>
<h2 id="references">References</h2>
<ul>
<li><a href="https://openstax.org/books/statistics/pages/4-2-mean-or-expected-value-and-standard-deviation">OpenStax: Mean or Expected Value and Standard Deviation</a></li>
<li><a href="https://plato.stanford.edu/entries/probability-interpret/">Stanford Encyclopedia of Philosophy: Interpretations of Probability</a></li>
<li><a href="https://plato.stanford.edu/entries/decision-theory/">Stanford Encyclopedia of Philosophy: Decision Theory</a></li>
<li><a href="https://plato.stanford.edu/entries/rationality-normative-utility/">Stanford Encyclopedia of Philosophy: Normative Theories of Rational Choice: Expected Utility</a></li>
<li><a href="https://www.ucl.ac.uk/anaesthesia/sites/anaesthesia/files/kahneman-tversky.pdf">Kahneman and Tversky: Prospect Theory: An Analysis of Decision under Risk</a></li>
</ul>
]]></content:encoded></item><item><title>Decision-Making: Judgment, Choice, and Commitment</title><link>https://moonment.net/en/notes/decision-making/</link><pubDate>Sat, 19 Sep 2026 15:30:00 +0800</pubDate><dc:creator>Moon</dc:creator><guid>https://moonment.net/en/notes/decision-making/</guid><description>Decision-making is not a moment of selection or a guarantee of good outcomes. It turns uncertain possibilities, evidence, and values into a revisable commitment to act.</description><content:encoded><![CDATA[<h2 id="what-is-decision-making">What Is Decision-Making?</h2>
<p>People perceive problems, form beliefs, imagine futures, and develop preferences. None of these activities by itself selects a course of action. A decision occurs when an agent resolves enough of the open possibilities for one direction to guide what happens next.</p>
<blockquote>
<p><strong>Decision-making is the process through which an agent responds to a practical situation by comparing possible actions, uncertain consequences, values, and constraints, then commits to a course of action.</strong></p>
</blockquote>
<p>The central transition is:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">several live possibilities
</span></span><span class="line"><span class="cl">→ judgment and trade-off
</span></span><span class="line"><span class="cl">→ one option gains practical priority
</span></span><span class="line"><span class="cl">→ planning and action
</span></span></code></pre></div><p>Commitment here is revisable. It means that the agent stops treating every possibility as equally open and begins allocating time, attention, authority, and resources. New evidence can reopen the decision.</p>
<h2 id="decision-judgment-choice-and-intention">Decision, Judgment, Choice, and Intention</h2>
<p>These terms often describe different parts of one episode.</p>
<table>
  <thead>
      <tr>
          <th>Concept</th>
          <th>Primary question</th>
          <th>Role</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Judgment</td>
          <td>What is true, likely, or important?</td>
          <td>Forms or revises belief</td>
      </tr>
      <tr>
          <td>Preference</td>
          <td>Which outcome do I favor?</td>
          <td>Orders outcomes or options</td>
      </tr>
      <tr>
          <td>Choice</td>
          <td>Which option was selected?</td>
          <td>Identifies the selected alternative</td>
      </tr>
      <tr>
          <td>Goal</td>
          <td>What state should be achieved?</td>
          <td>Specifies a desired result</td>
      </tr>
      <tr>
          <td>Intention</td>
          <td>What am I committed to doing?</td>
          <td>Organizes action across time</td>
      </tr>
      <tr>
          <td>Decision</td>
          <td>What course will govern this practical fork?</td>
          <td>Resolves alternatives into commitment</td>
      </tr>
      <tr>
          <td>Plan</td>
          <td>How will the course be carried out?</td>
          <td>Organizes steps, time, and resources</td>
      </tr>
      <tr>
          <td>Action</td>
          <td>What was actually done?</td>
          <td>Changes or attempts to change the world</td>
      </tr>
      <tr>
          <td>Outcome</td>
          <td>What eventually happened?</td>
          <td>Includes execution, environment, others, and luck</td>
      </tr>
  </tbody>
</table>
<p>“The project is likely to succeed” is a judgment. “Given its upside and our loss limit, we will fund the pilot” is a decision. The first does not entail the second. Action also requires values, constraints, alternatives, and an account of who bears the risk.</p>
<h2 id="the-structure-of-a-decision">The Structure of a Decision</h2>
<h3 id="a-practical-situation">A practical situation</h3>
<p>Why does a response seem necessary now? A decision problem begins with a conflict, opportunity, obstacle, or fork that matters to an agent.</p>
<h3 id="a-frame">A frame</h3>
<p>“Should we continue the project?”, “How can we reduce its failure risk?”, and “Which objective should we preserve?” frame the same situation differently. A frame determines which options and evidence become visible. Precise analysis cannot rescue the wrong problem.</p>
<h3 id="an-agent-and-authority">An agent and authority</h3>
<p>Who can make the selection effective? Who advises, who can veto, and who bears the consequences? Collective deliberation does not produce an operative decision unless an institution also defines authority and responsibility.</p>
<h3 id="ends-values-and-constraints">Ends, values, and constraints</h3>
<p>A goal specifies a desired state. Values explain why it matters. Constraints mark unacceptable means, risks, costs, or side effects.</p>
<p>Many hard decisions persist because several goods cannot be fully realized together. Revenue, safety, autonomy, fairness, speed, and loyalty may resist a common scale.</p>
<h3 id="feasible-options">Feasible options</h3>
<p>Success and failure are outcomes, not actions. Continue, stop, reduce scope, negotiate, run a pilot, or wait for information can be genuine options. Doing nothing and retaining the status quo also have consequences and should not disappear from the comparison.</p>
<h3 id="consequences-causation-and-uncertainty">Consequences, causation, and uncertainty</h3>
<p>A decision requires a view about what each action might change. This is a causal question, not merely an association. It also requires some representation of uncertainty, whether statistical, model-based, judgmental, or explicitly unknown.</p>
<p>Probability says how plausible an outcome is. It does not say how desirable, fair, or acceptable that outcome would be.</p>
<h3 id="a-decision-rule">A decision rule</h3>
<p>An agent might maximize expected value, limit ruin, protect a non-substitutable value, choose a robust option, preserve reversibility, or stop searching when an option clears an aspiration level. Different rules can select different actions. The rule itself therefore needs justification.</p>
<h3 id="commitment-execution-and-feedback">Commitment, execution, and feedback</h3>
<p>A decision must become a plan, allocation, instruction, or action. Observation then changes the next decision:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">decide
</span></span><span class="line"><span class="cl">→ execute
</span></span><span class="line"><span class="cl">→ observe
</span></span><span class="line"><span class="cl">→ compare expected and actual states
</span></span><span class="line"><span class="cl">→ revise beliefs, ends, or rules
</span></span><span class="line"><span class="cl">→ decide again
</span></span></code></pre></div><h2 id="three-questions-for-decision-theory">Three Questions for Decision Theory</h2>
<p>Decision research separates three projects.</p>
<h3 id="descriptive">Descriptive</h3>
<p>How do people actually decide? Psychology studies the effects of attention, memory, emotion, framing, defaults, social influence, and heuristics. The APA defines decision-making as the cognitive process of choosing between two or more alternatives. <a href="https://dictionary.apa.org/decision-making">APA Dictionary of Psychology</a></p>
<p>A recurring behavior does not become rational merely because it is common.</p>
<h3 id="normative">Normative</h3>
<p>How should coherent or rational choice be structured? Expected utility theory supplies one influential answer for choice under uncertainty:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">EU(A) = Σ P(Oᵢ | A) × U(Oᵢ)
</span></span></code></pre></div><p>It compares acts by weighting the utility of possible outcomes by their probabilities. Contemporary utility is often a representation of preference rather than a direct unit of money or happiness. <a href="https://plato.stanford.edu/entries/rationality-normative-utility/">Stanford Encyclopedia of Philosophy: Expected Utility</a></p>
<p>The formula does not generate the option set, validate causal assumptions, settle moral constraints, or identify whose preferences should count.</p>
<h3 id="prescriptive">Prescriptive</h3>
<p>How can a real person or organization improve a particular decision? Prescriptive work translates evidence and standards into usable practices:</p>
<ul>
<li>separate facts, estimates, values, and unknowns;</li>
<li>search for options suppressed by the initial frame;</li>
<li>use outside comparison classes;</li>
<li>specify stop, exit, and review conditions;</li>
<li>buy information only when it can change action;</li>
<li>test consequential assumptions through reversible steps;</li>
<li>record what was known before outcomes became visible.</li>
</ul>
<h2 id="bounded-rationality">Bounded Rationality</h2>
<p>No real agent has unlimited time, information, attention, or computation. Bounded rationality studies procedures that remain effective under those constraints. It does not simply label people irrational.</p>
<p>Herbert Simon&rsquo;s idea of satisficing replaces exhaustive optimization with a search process and an aspiration level: stop when an option is good enough relative to the costs of continuing. <a href="https://plato.stanford.edu/entries/bounded-rationality/">Stanford Encyclopedia of Philosophy: Bounded Rationality</a></p>
<p>The rational question can therefore be:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">Will the expected value of more information
</span></span><span class="line"><span class="cl">exceed the cost of search, delay, and lost opportunity?
</span></span></code></pre></div><p>Indefinite optimization can itself be a poor decision.</p>
<h2 id="behavioral-regularities-are-not-merely-noise">Behavioral Regularities Are Not Merely Noise</h2>
<p>Choices under risk depend on reference points, perceived gains and losses, presentation, and nonlinear sensitivity to probability. Prospect theory was developed to explain important patterns that standard economic models did not predict well. The 2002 Nobel Prize materials describe Daniel Kahneman&rsquo;s contribution as integrating psychological research on judgment and decision-making under uncertainty into economics. <a href="https://www.nobelprize.org/prizes/economic-sciences/2002/press-release/">Nobel Prize 2002</a></p>
<p>Calling a pattern a bias still requires a defensible benchmark. A shortcut can be poor under one environment and efficient under another once information and computation costs are included.</p>
<h2 id="why-decision-making-is-not-only-calculation">Why Decision-Making Is Not Only Calculation</h2>
<p>Facts constrain action but do not specify what should matter. A probability distribution cannot decide which losses are acceptable. A utility score can clarify a trade-off while concealing rights, identity, loyalty, or values that the agent refuses to exchange.</p>
<p>The presence of several nominal options also does not guarantee meaningful agency. Poverty, power, addiction, information control, and institutional defaults alter the feasible set and the conditions of responsibility.</p>
<p>Collective decisions add procedural values. Who had standing, access to evidence, voice, and veto power can matter independently of whether the final outcome was efficient.</p>
<h2 id="a-good-decision-can-have-a-bad-outcome">A Good Decision Can Have a Bad Outcome</h2>
<p>Four evaluations should remain separate:</p>
<table>
  <thead>
      <tr>
          <th>Evaluation</th>
          <th>Object</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Process quality</td>
          <td>Frame, options, evidence, and trade-offs</td>
      </tr>
      <tr>
          <td>Decision quality</td>
          <td>Defensibility of the commitment given information then available</td>
      </tr>
      <tr>
          <td>Execution quality</td>
          <td>Whether action implemented and adapted the decision</td>
      </tr>
      <tr>
          <td>Outcome quality</td>
          <td>Benefits, failures, and side effects that occurred</td>
      </tr>
  </tbody>
</table>
<p>A low-probability event can defeat a sound decision. Luck can rescue a careless one. Evaluating the original decision by information learned only afterward produces hindsight and outcome bias.</p>
<h2 id="a-practical-audit">A Practical Audit</h2>
<ol>
<li>What practical question actually requires resolution?</li>
<li>Has the initial frame excluded a better question?</li>
<li>Are inaction, delay, negotiation, or a reversible test real options?</li>
<li>Which statements are facts, causal assumptions, probabilities, values, or unknowns?</li>
<li>What mechanism connects each action to its expected effects?</li>
<li>What is being optimized, protected, or deliberately surrendered?</li>
<li>Can the worst plausible loss be borne?</li>
<li>Who decides, benefits, and bears risk?</li>
<li>Has the commitment entered plans, resources, and action?</li>
<li>Which new evidence would reopen the decision?</li>
</ol>
<p>Decision-making is the joint between understanding and action. It closes some possibilities so that agency can proceed, while preserving the capacity to learn from consequences.</p>
<blockquote>
<p><strong>A good decision does not guarantee a good result. It is a defensible, executable, accountable, and revisable commitment made from the evidence, values, and constraints available at the time.</strong></p>
</blockquote>
<h2 id="further-reading">Further Reading</h2>
<ul>
<li><a href="/en/notes/mental-models/">Mental Models: How We Represent, Predict, and Act</a></li>
<li><a href="/en/notes/human-thinking/">How Human Thinking Works: Representation, Reasoning, and Action</a></li>
</ul>
<h2 id="sources">Sources</h2>
<ul>
<li><a href="https://dictionary.apa.org/decision-making">APA Dictionary of Psychology: Decision Making</a></li>
<li><a href="https://plato.stanford.edu/entries/rationality-normative-utility/">Stanford Encyclopedia of Philosophy: Normative Theories of Rational Choice—Expected Utility</a></li>
<li><a href="https://plato.stanford.edu/entries/decision-theory-descriptive/">Stanford Encyclopedia of Philosophy: Descriptive Decision Theory</a></li>
<li><a href="https://plato.stanford.edu/entries/bounded-rationality/">Stanford Encyclopedia of Philosophy: Bounded Rationality</a></li>
<li><a href="https://www.nobelprize.org/prizes/economic-sciences/2002/press-release/">Nobel Prize 2002: Psychological and Experimental Economics</a></li>
</ul>
]]></content:encoded></item><item><title>Causation: Difference-Making, Mechanisms, and Reasons</title><link>https://moonment.net/en/notes/causality-causes-and-reasons/</link><pubDate>Wed, 16 Sep 2026 23:12:56 +0800</pubDate><dc:creator>Moon</dc:creator><guid>https://moonment.net/en/notes/causality-causes-and-reasons/</guid><description>A causal claim says more than one event followed or predicted another. It connects difference-making, intervention, counterfactual dependence, and mechanism while separating causes from evidence, reasons, and purposes.</description><content:encoded><![CDATA[<h2 id="what-does-a-causal-claim-say">What Does a Causal Claim Say?</h2>
<p>To call one thing a cause of another is to make a claim about how a difference is produced.</p>
<blockquote>
<p><strong>A factor is causally relevant to an outcome when its presence, absence, or variation makes a difference to how that outcome occurs, usually through conditions and processes that connect the two.</strong></p>
</blockquote>
<p>The factor may be an event, a standing condition, a behavior, a mental state, an institution, or a structural feature. The outcome may be a discrete event, but it may also be a change in magnitude, timing, form, or probability.</p>
<p>This definition is deliberately broader than determinism. Smoking can cause cancer without every smoker developing cancer. A treatment can cause recovery in the relevant population without curing every patient. Causes often alter the distribution of possible outcomes rather than fixing one outcome with certainty.</p>
<p>A minimal causal claim therefore has at least three parts:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">causal factor → connecting process → outcome
</span></span></code></pre></div><p>It also has a scope. A factor may be causal under one set of conditions and inert under another.</p>
<h2 id="causes-and-effects-are-roles-within-a-process">Causes and Effects Are Roles Within a Process</h2>
<p>An item is not permanently a cause or permanently an effect. Its role depends on which part of a process is under examination.</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">sleep deprivation
</span></span><span class="line"><span class="cl">→ impaired executive control
</span></span><span class="line"><span class="cl">→ more operational errors
</span></span><span class="line"><span class="cl">→ greater accident risk
</span></span></code></pre></div><p>Impaired executive control is an effect of sleep deprivation and a cause of later errors. An error can be an effect within one relation and a cause within the next.</p>
<p>This matters because causal language can create the illusion that the world has already been divided into two kinds of object. In practice, inquiry chooses an outcome and asks which earlier conditions and pathways help explain its production.</p>
<h2 id="sequence-association-and-prediction-are-not-yet-causation">Sequence, Association, and Prediction Are Not Yet Causation</h2>
<p>Three weaker relations are routinely mistaken for causal ones.</p>
<h3 id="temporal-sequence">Temporal sequence</h3>
<p>A cause normally precedes its effect, but precedence alone establishes very little. The rooster crows before sunrise; silencing the rooster does not delay the sun.</p>
<h3 id="statistical-association">Statistical association</h3>
<p>If conditioning on X changes the probability of Y,</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">P(Y | X) ≠ P(Y)
</span></span></code></pre></div><p>then X and Y are probabilistically dependent. This describes a distributional relation, not why it exists. Several structures remain possible:</p>
<table>
  <thead>
      <tr>
          <th>Structure</th>
          <th>Interpretation</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td><code>X → Y</code></td>
          <td>X causes Y</td>
      </tr>
      <tr>
          <td><code>X ← Y</code></td>
          <td>Y causes X</td>
      </tr>
      <tr>
          <td><code>X ← Z → Y</code></td>
          <td>Z is a common cause</td>
      </tr>
      <tr>
          <td><code>Xₜ → Yₜ → Xₜ₊₁</code></td>
          <td>X and Y form a feedback loop</td>
      </tr>
      <tr>
          <td>selection affects the sample</td>
          <td>the observed association is induced by who enters the data</td>
      </tr>
  </tbody>
</table>
<p>People who carry lighters may have a higher incidence of lung cancer. The lighter is not the relevant cause. Smoking helps explain both carrying the lighter and the increased risk.</p>
<p>Selection can induce an association that is absent in the wider population. Suppose severe illness and inadequate home care both increase the probability of hospitalization. Restricting a study to hospitalized patients conditions on their common effect. Within that selected sample, illness severity and home care can appear associated even when they were independent before selection. This is a form of collider or selection bias.</p>
<p>Small samples, repeated comparisons, changing measurement definitions, and correlated measurement errors can also produce unstable associations.</p>
<h3 id="prediction">Prediction</h3>
<p>A barometer can help predict a storm. Manipulating the needle does not change the weather. A predictor answers whether knowing X improves a forecast of Y. A causal variable answers whether changing X would change Y.</p>
<p>This distinction matters whenever a model is used for action. A system can predict accurately from proxies while offering no effective intervention. Predictive success does not by itself identify what should be changed.</p>
<p>Frequent customer-support contact may predict churn because product defects cause both help requests and departure:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">product defect → support contact
</span></span><span class="line"><span class="cl">product defect → churn
</span></span></code></pre></div><p>Removing the support channel would reduce the recorded predictor without repairing the defect. It could make churn worse. A predictor may be useful for finding risk while being the wrong target for intervention.</p>
<h2 id="most-outcomes-have-a-causal-architecture-not-a-single-cause">Most Outcomes Have a Causal Architecture, Not a Single Cause</h2>
<p>The demand for “the root cause” often compresses several explanatory tasks into one phrase.</p>
<p>A fire may depend on combustible material, oxygen, heat, building design, delayed detection, failed suppression, and organizational practices. One factor triggers ignition, another accelerates spread, another removes a barrier, and another explains why the dangerous configuration existed.</p>
<p>Useful distinctions include:</p>
<table>
  <thead>
      <tr>
          <th>Causal role</th>
          <th>Question</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Necessary condition</td>
          <td>Could the outcome occur without it?</td>
      </tr>
      <tr>
          <td>Sufficient condition</td>
          <td>Would it produce the outcome given the stated background?</td>
      </tr>
      <tr>
          <td>Contributing factor</td>
          <td>Does it raise the probability or severity?</td>
      </tr>
      <tr>
          <td>Trigger</td>
          <td>Does it initiate a prepared process?</td>
      </tr>
      <tr>
          <td>Background condition</td>
          <td>Does it enable other factors to operate?</td>
      </tr>
      <tr>
          <td>Sustaining cause</td>
          <td>Does it keep an existing outcome in place?</td>
      </tr>
      <tr>
          <td>Inhibitor</td>
          <td>Does it block or weaken a pathway?</td>
      </tr>
      <tr>
          <td>Structural cause</td>
          <td>Does it systematically shape many local conditions?</td>
      </tr>
  </tbody>
</table>
<p>Necessary and sufficient conditions should not be confused. Oxygen is necessary for ordinary combustion, but oxygen alone is not sufficient for a building fire.</p>
<p>Many causes are components of a larger sufficient package. Different packages may also produce the same outcome. This is why removing one factor may fail to prevent an effect even when that factor was causally active: another sufficient pathway may remain.</p>
<h2 id="causal-relations-form-chains-forks-and-feedback-loops">Causal Relations Form Chains, Forks, and Feedback Loops</h2>
<p>Several basic patterns recur.</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">multiple causes:      X₁ + X₂ + X₃ → Y
</span></span><span class="line"><span class="cl">multiple effects:     X → Y₁, Y₂, Y₃
</span></span><span class="line"><span class="cl">mediation:            X → M₁ → M₂ → Y
</span></span><span class="line"><span class="cl">common cause:         X ← Z → Y
</span></span><span class="line"><span class="cl">feedback over time:   Xₜ → Yₜ → Xₜ₊₁
</span></span></code></pre></div><p>The time index is essential in feedback systems. Stress may produce insomnia, which produces more stress on the following day. Popularity may generate reviews, and those reviews may become social proof that produces later popularity.</p>
<p>Without time, the relation looks circular. With time, it becomes a sequence of reciprocal effects.</p>
<h2 id="probabilistic-causation-does-not-mean-mere-correlation">Probabilistic Causation Does Not Mean Mere Correlation</h2>
<p>A first approximation says that a cause raises the probability of its effect:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">P(Y | X) &gt; P(Y | not-X)
</span></span></code></pre></div><p>But this remains an observational comparison. If X is more common among people who differ in other relevant ways, the inequality may reflect confounding rather than an effect of X.</p>
<p>The causal question is closer to:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">P(Y | do(X)) &gt; P(Y | do(not-X))
</span></span></code></pre></div><p>The <code>do</code> operator represents setting X by intervention while preserving the rest of the causal model as specified. It distinguishes observing X from changing X. That distinction is central to structural approaches to causal inference associated with Judea Pearl. <a href="https://ftp.cs.ucla.edu/pub/stat_ser/ACMBook-published-2022.pdf">Probabilistic and Causal Inference: The Works of Judea Pearl</a></p>
<p>A probabilistic cause can be real even when:</p>
<ul>
<li>the effect does not occur in a particular case;</li>
<li>the effect sometimes occurs without that cause;</li>
<li>the same cause has different effects in different contexts;</li>
<li>several causal pathways compete.</li>
</ul>
<p>Individual outcomes and population effects are different claims. One patient recovering without treatment does not show that the treatment has no causal effect. One treated patient failing to recover does not show that the treatment is ineffective in the relevant population.</p>
<h2 id="causation-and-causal-inference-are-different-problems">Causation and Causal Inference Are Different Problems</h2>
<p>Causation concerns relations in the world:</p>
<blockquote>
<p>What actually contributed to the production of the outcome?</p>
</blockquote>
<p>Causal inference concerns our epistemic position:</p>
<blockquote>
<p>What justifies believing that a particular factor was causal?</p>
</blockquote>
<p>The distinction is basic:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">causal relation ≠ method for discovering a causal relation
</span></span></code></pre></div><p>A causal relation can exist before anyone recognizes it. An elegant explanation can be accepted even though its causal structure is wrong.</p>
<p>Inquiry usually begins with an observed effect. It then works backward to candidate causes and forward again to predictions. This requires several forms of reasoning.</p>
<h2 id="how-are-causes-inferred">How Are Causes Inferred?</h2>
<h3 id="abduction-generates-candidate-explanations">Abduction generates candidate explanations</h3>
<p>Wet pavement may be explained by rain, a broken pipe, a street-cleaning vehicle, or deliberate watering. Abduction asks which hypothesis would best explain the observation.</p>
<p>Abduction is ampliative: its conclusion goes beyond what is logically contained in the evidence. Its appeal to explanatory considerations distinguishes it from induction based primarily on observed frequencies. <a href="https://plato.stanford.edu/entries/abduction/">Stanford Encyclopedia of Philosophy: Abduction</a></p>
<p>Generating a good explanation is not the same as proving it.</p>
<h3 id="deduction-derives-consequences">Deduction derives consequences</h3>
<p>If a pipe is broken, water should continue under specified weather conditions, concentrate near the line, and respond to a closed valve. Deduction turns a hypothesis into testable expectations.</p>
<p>Failure of a prediction can weaken the hypothesis. Success does not uniquely confirm it when rival hypotheses predict the same evidence.</p>
<h3 id="induction-extends-patterns">Induction extends patterns</h3>
<p>Repeated observations can support a generalization. The inference remains vulnerable to unrepresentative samples, environmental change, hidden common causes, and selective observation.</p>
<h3 id="rival-explanations-must-be-compared">Rival explanations must be compared</h3>
<p>The strongest evidence is often not evidence that fits one hypothesis. It is evidence that one hypothesis predicts and its competitors do not.</p>
<h3 id="bayesian-updating-revises-confidence">Bayesian updating revises confidence</h3>
<p>For a hypothesis H and evidence E:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">P(H | E) = P(E | H)P(H) / P(E)
</span></span></code></pre></div><p>Bayes&rsquo; rule disciplines how confidence changes when evidence arrives. It does not identify the causal graph by itself. Priors, likelihoods, variable choices, and the hypothesis space all depend on substantive assumptions.</p>
<p>A fuller cycle is:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">observe an outcome
</span></span><span class="line"><span class="cl">→ generate candidate causes
</span></span><span class="line"><span class="cl">→ derive discriminating predictions
</span></span><span class="line"><span class="cl">→ collect evidence or intervene
</span></span><span class="line"><span class="cl">→ update confidence
</span></span><span class="line"><span class="cl">→ revise the causal model
</span></span></code></pre></div><h2 id="counterfactuals-ask-what-would-have-happened-otherwise">Counterfactuals Ask What Would Have Happened Otherwise</h2>
<p>Counterfactual analysis asks:</p>
<blockquote>
<p>If X had not occurred, would Y still have occurred?</p>
</blockquote>
<p>This captures a central intuition: causes make a difference. But the unobserved alternative creates the fundamental problem of causal inference. The same patient cannot both receive and not receive a treatment at the same moment. The same firm cannot simultaneously adopt and reject the same strategy under identical conditions.</p>
<p>Randomized trials, matched comparisons, natural experiments, historical controls, and causal models are different ways of estimating the missing alternative.</p>
<p>Simple counterfactual dependence is not a complete theory. If two independent fires were each sufficient to destroy a building, removing one would not prevent the destruction. Preemption and overdetermination show why actual causation can be more complex than a single but-for test. <a href="https://plato.stanford.edu/entries/causation-counterfactual/">Stanford Encyclopedia of Philosophy: Counterfactual Theories of Causation</a></p>
<h2 id="interventions-ask-what-changing-a-variable-would-do">Interventions Ask What Changing a Variable Would Do</h2>
<p>Observational questions compare naturally occurring groups. Interventional questions ask what would happen if a variable were deliberately set.</p>
<p>Suppose patients receiving a treatment are sicker on average. The treatment may appear associated with worse outcomes because severity influenced who received it. Random allocation can reduce this source of confounding by making groups comparable in expectation.</p>
<p>Experiments are not automatically decisive. Attrition, noncompliance, measurement error, short follow-up, small samples, and differences between the study setting and the target setting can all limit a conclusion.</p>
<p>Observational studies can still support causal claims when their design, assumptions, controls, and sensitivity analyses address the relevant alternatives. The real question is how well the design separates the proposed effect from competing explanations.</p>
<h2 id="mechanisms-explain-how-the-difference-is-produced">Mechanisms Explain How the Difference Is Produced</h2>
<p>Mechanistic inquiry opens the arrow:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">X → M₁ → M₂ → Y
</span></span></code></pre></div><p>For example:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">chronic sleep loss
</span></span><span class="line"><span class="cl">→ impaired executive function
</span></span><span class="line"><span class="cl">→ weaker attentional control
</span></span><span class="line"><span class="cl">→ more errors
</span></span></code></pre></div><p>Mechanisms can identify intervention points, explain variation across contexts, and show why a relationship should generalize. They also reveal mediators that should not be treated as independent background variables.</p>
<p>A plausible mechanism is not enough. Post hoc stories are easy to invent. The intermediate stages need independent evidence.</p>
<p>Counterfactual, interventional, and mechanistic approaches answer different questions:</p>
<table>
  <thead>
      <tr>
          <th>Approach</th>
          <th>Central question</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Counterfactual</td>
          <td>What would happen without X?</td>
      </tr>
      <tr>
          <td>Intervention</td>
          <td>What would happen if X were changed?</td>
      </tr>
      <tr>
          <td>Mechanism</td>
          <td>Through what process does X affect Y?</td>
      </tr>
  </tbody>
</table>
<p>The approaches reinforce one another without becoming interchangeable.</p>
<h2 id="the-direction-of-explanation-can-oppose-the-direction-of-causation">The Direction of Explanation Can Oppose the Direction of Causation</h2>
<p>The causal direction may be:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">rain → wet pavement
</span></span></code></pre></div><p>The direction of inference may be:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">wet pavement → evidence for recent rain
</span></span></code></pre></div><p>The pavement does not cause the earlier rain. The effect supplies evidence about its possible cause.</p>
<p>Medicine, engineering diagnosis, historical inquiry, and accident investigation routinely reason from traces to causes. The mistake is not beginning with an effect. The mistake is treating an explanation that fits the effect as a cause already established.</p>
<h2 id="when-an-effect-is-mistaken-for-a-cause">When an Effect Is Mistaken for a Cause</h2>
<p>Several distinct errors are often grouped together.</p>
<h3 id="reverse-causation">Reverse causation</h3>
<p>Severe illness may increase treatment uptake. A raw association between treatment and severity can then be misread as evidence that treatment caused the severity.</p>
<h3 id="feedback">Feedback</h3>
<p>An effect at one stage can become a cause at the next. Initial popularity produces reviews; reviews produce social proof; social proof contributes to later popularity. This is a real causal loop over time, not simply a mistaken direction.</p>
<h3 id="selection-on-successful-cases">Selection on successful cases</h3>
<p>If successful people often wake early, early rising may be a cause, an effect of their circumstances, a correlate of other traits, or a minor contributor. The unsuccessful early risers omitted from the sample matter.</p>
<h3 id="hindsight-and-outcome-bias">Hindsight and outcome bias</h3>
<p>After a success, risk-taking is called vision. After a failure, the same behavior is called recklessness. Knowing the outcome changes the story told about the decision.</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">quality of a decision ≠ quality of its realized outcome
</span></span></code></pre></div><p>A decision should be assessed using the information, probabilities, aims, and constraints available when it was made.</p>
<h2 id="causes-evidence-and-reasons-answer-different-questions">Causes, Evidence, and Reasons Answer Different Questions</h2>
<p>The sentence “because the pavement is wet, it probably rained” cites evidence. “Because it rained, the pavement became wet” cites a cause.</p>
<p>Human action introduces further distinctions.</p>
<h3 id="motivating-reasons">Motivating reasons</h3>
<p>A person may resign because she believes continued work is harming her health. The consideration under which she acts is her motivating reason.</p>
<h3 id="normative-reasons">Normative reasons</h3>
<p>Actual harm to health may count in favor of resigning. A normative reason concerns what supports or justifies an action, whether or not the agent acted for it.</p>
<h3 id="explanatory-reasons">Explanatory reasons</h3>
<p>Exhaustion, fear, or resentment may explain an action without justifying it.</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">what explains an action ≠ what justifies it
</span></span></code></pre></div><h3 id="stated-reasons">Stated reasons</h3>
<p>What an agent says afterward may be an accurate report, a partial account, a socially acceptable presentation, or a rationalization.</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">stated reason ≠ actual motivation ≠ normative justification
</span></span></code></pre></div><p>Contemporary philosophy of action commonly distinguishes normative, motivating, and explanatory reasons according to whether they favor, guide, or explain action. <a href="https://plato.stanford.edu/entries/reasons-just-vs-expl/">Stanford Encyclopedia of Philosophy: Reasons for Action</a></p>
<p>Donald Davidson argued that explaining an intentional action by the agent&rsquo;s reasons can also be causal explanation. A relevant belief and desire do not merely make an action intelligible; when they actually produce it, they are among its causes. <a href="https://plato.stanford.edu/entries/davidson/">Stanford Encyclopedia of Philosophy: Donald Davidson</a></p>
<h2 id="purposes-are-represented-in-the-present">Purposes Are Represented in the Present</h2>
<p>“She exercises in order to become healthier” can sound as if a future outcome causes a present action. The future state does not reach backward in time.</p>
<p>The operative causal structure is present:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">desire for better health
</span></span><span class="line"><span class="cl">+ belief that exercise will help
</span></span><span class="line"><span class="cl">+ intention to exercise
</span></span><span class="line"><span class="cl">→ present action
</span></span></code></pre></div><p>Purpose concerns the outcome an agent seeks. Expectation concerns what the agent believes will occur. Intention organizes present and future action. The actual outcome is what later happens. These can diverge.</p>
<p>A self-fulfilling expectation follows the same pattern. The future outcome is not its own earlier cause. A present expectation changes behavior, and the changed behavior helps produce the expected outcome.</p>
<h2 id="why-philosophers-disagree-about-causation">Why Philosophers Disagree About Causation</h2>
<p>Different theories emphasize different parts of the concept.</p>
<p>Aristotle&rsquo;s four causes addressed material, form, source of change, and end. His notion of <em>aitia</em> was broader than the modern search for efficient production. It organized several kinds of answer to a why-question.</p>
<p>Hume challenged the idea that necessary connection is directly perceived. Experience presents succession and repeated conjunction; the necessity attributed to the sequence requires further explanation.</p>
<p>Kant treated causal ordering as a condition for objective experience. A sequence of perceptions must be distinguished from a perception of an objective sequence of events.</p>
<p>Modern families of theory isolate different features:</p>
<table>
  <thead>
      <tr>
          <th>Family</th>
          <th>Emphasis</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Regularity theories</td>
          <td>stable patterns between cause and effect</td>
      </tr>
      <tr>
          <td>Probabilistic theories</td>
          <td>changes in the probability of an outcome</td>
      </tr>
      <tr>
          <td>Counterfactual theories</td>
          <td>what would differ without the cause</td>
      </tr>
      <tr>
          <td>Mechanistic theories</td>
          <td>processes that transmit causal influence</td>
      </tr>
      <tr>
          <td>Interventionist theories</td>
          <td>systematic changes under manipulation</td>
      </tr>
      <tr>
          <td>Structural causal models</td>
          <td>variables, equations, graphs, and counterfactual states</td>
      </tr>
  </tbody>
</table>
<p>No single entry in the table should be treated as the whole meaning of causation in every domain. Together they explain why causal judgment involves regularity, difference-making, production, and control.</p>
<h2 id="a-practical-causal-analysis">A Practical Causal Analysis</h2>
<p>When asked why something happened, proceed in this order:</p>
<ol>
<li><strong>Specify the outcome.</strong> Replace broad labels with an observable event, state, or measure.</li>
<li><strong>Add time.</strong> Mark when candidate causes, intermediate stages, and outcomes occurred.</li>
<li><strong>List rival structures.</strong> Include reverse causation, common causes, selection, and feedback.</li>
<li><strong>Draw the pathways.</strong> Identify confounders, mediators, inhibitors, and alternative routes.</li>
<li><strong>Ask the counterfactual.</strong> What would probably happen without the factor?</li>
<li><strong>Seek a comparison or intervention.</strong> What design could reveal the difference made by changing it?</li>
<li><strong>Test the mechanism.</strong> What intermediate evidence should exist if the account is correct?</li>
<li><strong>Separate causes from reasons.</strong> Is the claim about production, evidence, motivation, or justification?</li>
<li><strong>State assumptions and scope.</strong> Which conclusions depend on which model and population?</li>
<li><strong>Keep uncertainty visible.</strong> Distinguish a supported causal claim from an unresolved hypothesis.</li>
</ol>
<p>A responsible conclusion may therefore read:</p>
<blockquote>
<p>Under the stated assumptions and current evidence, X probably raises the risk of Y through mechanism M. Z remains a plausible source of residual confounding.</p>
</blockquote>
<p>That is not evasive language. It specifies what is known, why it is believed, and where the inference can fail.</p>
<h2 id="a-product-experiment-across-logic-probability-and-causation">A Product Experiment Across Logic, Probability, and Causation</h2>
<p>Suppose a product team asks whether push reminders increase task completion.</p>
<p>The logical and operational layer comes first. A task counts as complete only when a completion timestamp exists. Eligible users must have an unfinished qualifying task before the reminder is assigned. Without consistent events, denominators, and time windows, the comparison is not well formed.</p>
<p>An observational analysis may then find:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">completion among reminded users: 60%
</span></span><span class="line"><span class="cl">completion among non-reminded users: 40%
</span></span></code></pre></div><p>The probability difference is real in the data but does not yet identify an effect. More active users may enable notifications and complete tasks more often for independent reasons.</p>
<p>A causal design can randomly assign eligible users to one reminder or no reminder, preserve the same outcome definition and observation window, and compare completion rates. Randomization aims to prevent known and unknown background differences from being systematically concentrated in one group.</p>
<p>The proposed mechanism should also leave intermediate traces:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">reminder
</span></span><span class="line"><span class="cl">→ attention
</span></span><span class="line"><span class="cl">→ app open
</span></span><span class="line"><span class="cl">→ task view
</span></span><span class="line"><span class="cl">→ completion
</span></span></code></pre></div><p>An increase in completion without corresponding evidence along the path calls for rival explanations. Even a positive average treatment effect has a scope: duration, reminder frequency, subgroups, timing shifts, notification opt-outs, and annoyance may all matter.</p>
<p>A bounded conclusion would therefore say:</p>
<blockquote>
<p>For eligible users in this experiment, one randomly assigned reminder increased completion by the estimated amount within the stated observation window. Generalization to other users, frequencies, and time horizons remains to be tested.</p>
</blockquote>
<h2 id="conclusion">Conclusion</h2>
<p>Causation concerns how a difference in one part of the world helps produce a difference in another.</p>
<p>It is not identical to sequence, association, prediction, a persuasive narrative, an agent&rsquo;s stated reason, or a retrospective judgment. A serious causal account connects several kinds of support:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">temporal order
</span></span><span class="line"><span class="cl">+ probabilistic difference
</span></span><span class="line"><span class="cl">+ counterfactual comparison
</span></span><span class="line"><span class="cl">+ intervention
</span></span><span class="line"><span class="cl">+ mechanism
</span></span><span class="line"><span class="cl">+ comparison with rival explanations
</span></span></code></pre></div><p>Causal inference then adds the methods by which such an account is discovered and revised:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">abduction generates hypotheses
</span></span><span class="line"><span class="cl">+ deduction derives predictions
</span></span><span class="line"><span class="cl">+ induction extends patterns
</span></span><span class="line"><span class="cl">+ experiments and comparisons test differences
</span></span><span class="line"><span class="cl">+ Bayesian updating revises confidence
</span></span><span class="line"><span class="cl">+ mechanistic inquiry opens the causal pathway
</span></span></code></pre></div><p>The aim is not to attach one definitive “root cause” to every outcome. It is to build a causal model that can be criticized, tested, used for intervention, and revised when new evidence arrives.</p>
]]></content:encoded></item><item><title>Probability and Bayes: Uncertainty, Evidence, and Rational Updating</title><link>https://moonment.net/en/notes/probability-and-bayes/</link><pubDate>Sat, 12 Sep 2026 22:03:10 +0800</pubDate><dc:creator>Moon</dc:creator><guid>https://moonment.net/en/notes/probability-and-bayes/</guid><description>The mathematics of probability is precise, but its meaning is contested. Bayes' rule shows how conditional probabilities relate without turning uncertainty into truth.</description><content:encoded><![CDATA[<p>Probability has an unusual philosophical structure. Its mathematics can be exact even when people disagree about what the number represents.</p>
<p>When someone says that an event has probability 0.7, they might be describing a long-run frequency, a physical tendency, the support supplied by evidence, or their own rational degree of confidence. The same formal rules can operate across these interpretations. The rules tell us how probabilities must fit together; they do not by themselves tell us what probabilities are.</p>
<p>Bayes’ rule works inside that formal structure. It relates an initial probability, new evidence, and an updated probability. It is indispensable for reasoning under uncertainty, but it does not guarantee that the starting assumptions, evidence, or model are correct.</p>
<p>This essay asks what a probability means and how Bayesian updating changes a judgment. <a href="/en/notes/logic-and-probability/">Logic and Probability</a> compares evidential support with entailment; <a href="/en/notes/what-is-logic/">Logic</a> examines validity within inference itself.</p>
<h2 id="one-calculus-several-meanings">One calculus, several meanings</h2>
<p>Consider four claims:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">A fair coin has a 50% probability of landing heads.
</span></span><span class="line"><span class="cl">A radium atom has a certain probability of decaying within a year.
</span></span><span class="line"><span class="cl">The evidence gives the defendant a 20% probability of guilt.
</span></span><span class="line"><span class="cl">I am 80% confident that the train will arrive on time.
</span></span></code></pre></div><p>They all use probability language, but they do not obviously describe the same kind of property.</p>
<p>The first may be grounded in symmetry. The second appears to describe a physical process. The third concerns how evidence bears on a hypothesis. The fourth describes a person’s graded confidence.</p>
<p>The philosophy of probability asks what makes statements like these true or reasonable. The Stanford Encyclopedia of Philosophy groups the central possibilities around physical probability, evidential support, and degrees of confidence, while noting that their boundaries can overlap. <a href="https://plato.stanford.edu/entries/probability-interpret/">Stanford Encyclopedia of Philosophy: Interpretations of Probability</a></p>
<h2 id="the-axioms-constrain-probability-without-interpreting-it">The axioms constrain probability without interpreting it</h2>
<p>A probability model begins with:</p>
<ul>
<li>a sample space <code>Ω</code>, containing possible outcomes;</li>
<li>events represented as subsets of <code>Ω</code>;</li>
<li>a function <code>P</code> assigning numbers to those events.</li>
</ul>
<p>The standard axioms require that probabilities are nonnegative, that <code>P(Ω) = 1</code>, and that mutually exclusive events add appropriately.</p>
<p>For an ordinary six-sided die:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">Ω = {1, 2, 3, 4, 5, 6}
</span></span><span class="line"><span class="cl">P(even) = P({2, 4, 6}) = 1/2
</span></span></code></pre></div><p>These rules establish coherence. They do not tell us why each face receives probability <code>1/6</code>. That assignment might rely on physical symmetry, observed frequencies, a model of the die, or a state of information.</p>
<p>This separation is crucial:</p>
<blockquote>
<p><strong>Probability theory specifies relations among probability values. An interpretation explains what those values mean and how they should be assigned.</strong></p>
</blockquote>
<h2 id="chance-in-the-world-and-uncertainty-in-knowledge">Chance in the world and uncertainty in knowledge</h2>
<p>Some uncertainty appears to concern how the world behaves. Before a coin lands, the physical process may be modeled as chancy. Radioactive decay is commonly represented probabilistically even when the experimental conditions are carefully controlled.</p>
<p>Other uncertainty is plainly epistemic. A coin may already have landed under a cup. The result is fixed, but an observer who cannot see it may assign equal confidence to heads and tails.</p>
<p>This produces a persistent question:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">Is probability a feature of the world,
</span></span><span class="line"><span class="cl">a relation between evidence and propositions,
</span></span><span class="line"><span class="cl">or a feature of an agent&#39;s information?
</span></span></code></pre></div><p>Frequency interpretations connect probability with proportions in repeated trials. Propensity accounts treat probability as a tendency or disposition of a setup to produce outcomes. Logical or evidential accounts connect probability with how strongly evidence supports a conclusion. Subjective or personalist accounts represent an agent’s coherent degree of belief.</p>
<p>No interpretation is automatically best for every use. A weather forecast, a quantum transition, a clinical trial, and a person’s confidence in a business decision may require different explanatory work even when they use the same calculus.</p>
<h2 id="conditional-probability-makes-information-explicit">Conditional probability makes information explicit</h2>
<p>Probabilities are often relative to conditions:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">P(A | B)
</span></span></code></pre></div><p>reads as the probability of <code>A</code> given <code>B</code>. When <code>P(B) &gt; 0</code>:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">P(A | B) = P(A ∩ B) / P(B)
</span></span></code></pre></div><p>Suppose 40 of 100 people carry umbrellas, and 30 of those 40 arrive wet. Then:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">P(wet | umbrella) = 30 / 40 = 0.75
</span></span></code></pre></div><p>This is not necessarily the same as:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">P(umbrella | wet)
</span></span></code></pre></div><p>Reversing the condition changes the reference class. Much faulty probabilistic reasoning comes from treating these two quantities as interchangeable.</p>
<h2 id="bayes-rule-reverses-a-conditional-relation">Bayes’ rule reverses a conditional relation</h2>
<p>From the definition of conditional probability:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">P(H ∩ E) = P(E | H)P(H)
</span></span><span class="line"><span class="cl">P(H ∩ E) = P(H | E)P(E)
</span></span></code></pre></div><p>Therefore:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">P(H | E) = P(E | H)P(H) / P(E)
</span></span></code></pre></div><p>Here:</p>
<ul>
<li><code>H</code> is a hypothesis;</li>
<li><code>E</code> is observed evidence;</li>
<li><code>P(H)</code> is the prior probability;</li>
<li><code>P(E | H)</code> is the likelihood of the evidence under the hypothesis;</li>
<li><code>P(H | E)</code> is the posterior probability.</li>
</ul>
<p>The denominator can be expanded across competing hypotheses. If <code>H</code> and <code>¬H</code> exhaust the possibilities:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">P(E) = P(E | H)P(H) + P(E | ¬H)P(¬H)
</span></span></code></pre></div><p>Bayes’ rule is an identity. The controversy begins when it is used as a general account of learning: which priors are rational, which hypotheses belong in the model, and what should count as evidence?</p>
<h2 id="why-base-rates-change-the-meaning-of-a-positive-test">Why base rates change the meaning of a positive test</h2>
<p>Suppose a condition affects 1% of a population. A test has:</p>
<ul>
<li>90% sensitivity: <code>P(positive | condition) = 0.90</code>;</li>
<li>5% false-positive rate: <code>P(positive | no condition) = 0.05</code>.</li>
</ul>
<p>The probability of the condition after a positive result is:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">P(condition | positive)
</span></span><span class="line"><span class="cl">= (0.90 × 0.01) / [(0.90 × 0.01) + (0.05 × 0.99)]
</span></span><span class="line"><span class="cl">= 0.009 / 0.0585
</span></span><span class="line"><span class="cl">≈ 0.154
</span></span></code></pre></div><p>So the posterior probability is about 15.4%, not 90%.</p>
<p>A frequency representation makes this intuitive. Among 10,000 people:</p>
<ul>
<li>about 100 have the condition, and 90 test positive;</li>
<li>about 9,900 do not, and roughly 495 test positive;</li>
<li>among 585 positive results, only 90 are true positives.</li>
</ul>
<p>The test is informative: probability rises from 1% to about 15.4%. But sensitivity answers how often the test detects the condition when it is present. It does not directly answer how often the condition is present when the test is positive. <a href="https://online.stat.psu.edu/stat414/Lesson06">Penn State STAT 414: Bayes’ Theorem</a></p>
<h2 id="bayesianism-turns-updating-into-a-norm-of-belief">Bayesianism turns updating into a norm of belief</h2>
<p>Bayesian epistemology represents confidence as a <strong>credence</strong>, a value between 0 and 1. It then asks how credences ought to fit together and how they ought to change when evidence arrives.</p>
<p>This approach separates two norms:</p>
<ol>
<li><strong>Synchronic coherence:</strong> probabilities held at one time should satisfy the probability axioms.</li>
<li><strong>Diachronic updating:</strong> beliefs should change in a disciplined way as evidence changes.</li>
</ol>
<p>Bayesian conditionalization is one proposed updating rule. If an agent becomes certain of evidence <code>E</code>, the new credence in <code>H</code> should often equal the old conditional credence <code>P(H | E)</code>.</p>
<p>This gives a precise account of graded belief. It also explains why rationality need not require certainty. Two people can assign different probabilities while both responding coherently to evidence, especially when their prior information differs. <a href="https://plato.stanford.edu/entries/epistemology-bayesian/">Stanford Encyclopedia of Philosophy: Bayesian Epistemology</a></p>
<h2 id="updating-cannot-repair-a-bad-model-by-itself">Updating cannot repair a bad model by itself</h2>
<p>Bayesian reasoning is only as good as the space within which it updates.</p>
<h3 id="priors-require-justification">Priors require justification</h3>
<p>Some priors follow from measured base rates or well-tested models. Others express limited information, expert judgment, convention, or convenience. Labeling a number “prior” does not make it objective.</p>
<h3 id="the-hypothesis-space-may-omit-the-truth">The hypothesis space may omit the truth</h3>
<p>If a diagnosis system considers only three diseases while the patient has a fourth, all posterior probability will be redistributed among the wrong options.</p>
<h3 id="likelihoods-depend-on-a-model">Likelihoods depend on a model</h3>
<p><code>P(E | H)</code> assumes a relation between hypothesis and evidence. Measurement error, selection bias, dependence among observations, and changing environments can make the assumed likelihood unreliable.</p>
<h3 id="zero-priors-can-block-learning">Zero priors can block learning</h3>
<p>If <code>P(H) = 0</code>, ordinary Bayesian updating leaves the posterior at zero regardless of the evidence. Absolute certainty assigned too early can make a model unable to recover.</p>
<h3 id="evidence-selection-is-not-neutral">Evidence selection is not neutral</h3>
<p>Someone must decide what to measure, which observations are relevant, and how the data are encoded. Updating can be formally correct while the evidence pipeline remains biased or incomplete.</p>
<h2 id="a-posterior-probability-is-not-truth">A posterior probability is not truth</h2>
<p>A posterior expresses probability under a model and body of evidence. It is not a direct transformation of uncertainty into fact.</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">high posterior probability
</span></span><span class="line"><span class="cl">≠ verified truth
</span></span><span class="line"><span class="cl">≠ complete hypothesis space
</span></span><span class="line"><span class="cl">≠ reliable measurement
</span></span><span class="line"><span class="cl">≠ justified decision in every context
</span></span></code></pre></div><p>Decisions also depend on consequences. A 5% probability may justify action when the possible harm is catastrophic; a 95% probability may still be insufficient for an irreversible accusation.</p>
<p>Probability helps make uncertainty explicit. Bayes’ rule makes one important form of learning explicit. Their philosophical value lies in disciplined revision, not in a promise of certainty.</p>
<h2 id="sources">Sources</h2>
<ul>
<li><a href="https://plato.stanford.edu/entries/probability-interpret/">Stanford Encyclopedia of Philosophy: Interpretations of Probability</a></li>
<li><a href="https://plato.stanford.edu/entries/epistemology-bayesian/">Stanford Encyclopedia of Philosophy: Bayesian Epistemology</a></li>
<li><a href="https://online.stat.psu.edu/stat414/Lesson06">Penn State STAT 414: Bayes’ Theorem</a></li>
</ul>
]]></content:encoded></item></channel></rss>