<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/"><channel><title>Ai-Agent on Moonment</title><link>https://moonment.net/en/tags/ai-agent/</link><description>Moon's notes on concepts, real projects, and reasoning open to review.</description><generator>Hugo</generator><language>en-US</language><managingEditor>Moon</managingEditor><webMaster>Moon</webMaster><copyright>© 2026 Moonment</copyright><lastBuildDate>Tue, 29 Sep 2026 13:55:00 +0800</lastBuildDate><atom:link href="https://moonment.net/en/tags/ai-agent/index.xml" rel="self" type="application/rss+xml"/><item><title>Systems: Boundaries, Interactions, and AI</title><link>https://moonment.net/en/notes/what-is-a-system/</link><pubDate>Sun, 27 Sep 2026 21:07:57 +0800</pubDate><dc:creator>Moon</dc:creator><guid>https://moonment.net/en/notes/what-is-a-system/</guid><description>A system is more than a collection of parts. This essay explains boundaries, interactions, state, feedback, emergence, purpose, and why an AI model is only one component of an AI system.</description><content:encoded><![CDATA[<p>A pile of parts is not yet a system. A list of employees is not yet an organization. A language model is not, by itself, the complete application that a user encounters.</p>
<p>The word <em>system</em> becomes useful when parts are related in ways that produce persistent behavior at the level of a whole.</p>
<blockquote>
<p><strong>A system is a bounded set of interdependent elements whose organization and interactions produce behavior over time within an environment.</strong></p>
</blockquote>
<p>This definition contains several commitments. A system has elements, but it cannot be understood from an inventory alone. It has a boundary, though that boundary depends partly on the question being asked. It has a state that can change. It interacts with an environment. Its overall behavior depends on relations among parts, not merely on the parts considered separately.</p>
<h2 id="the-word-and-its-central-idea">The word and its central idea</h2>
<p>English <em>system</em> comes through Latin <em>systema</em> from Greek <em>systēma</em>, an organized whole composed of parts. Its roots carry the idea of things standing together rather than existing as an unrelated assortment.<a href="https://www.etymonline.com/word/system">Etymonline: system</a></p>
<p>The term now covers very different objects:</p>
<ul>
<li>the solar system;</li>
<li>a nervous system;</li>
<li>an ecosystem;</li>
<li>a legal system;</li>
<li>an organization;</li>
<li>a payment system;</li>
<li>a software system;</li>
<li>an AI system.</li>
</ul>
<p>These examples do not share one material composition or one kind of purpose. What they share is an analytical form: distinguishable elements participate in relations that sustain some pattern of behavior at the level of a whole.</p>
<h2 id="what-must-a-system-contain">What must a system contain?</h2>
<p>The smallest useful account of a system normally identifies:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">elements
</span></span><span class="line"><span class="cl">+ relations
</span></span><span class="line"><span class="cl">+ a boundary
</span></span></code></pre></div><p>An account of how the system operates also needs:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">state
</span></span><span class="line"><span class="cl">+ rules or mechanisms of change
</span></span><span class="line"><span class="cl">+ an environment
</span></span><span class="line"><span class="cl">+ inputs and outputs
</span></span><span class="line"><span class="cl">+ time
</span></span></code></pre></div><p>For an engineered system, further questions become central:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">intended purpose
</span></span><span class="line"><span class="cl">+ constraints
</span></span><span class="line"><span class="cl">+ performance criteria
</span></span><span class="line"><span class="cl">+ authority and responsibility
</span></span></code></pre></div><p>Engineering standards commonly define a system as interacting elements organized to achieve one or more stated purposes.<a href="https://www.iso.org/obp/ui?_escaped_fragment_=iso%3Astd%3Aiso-iec-ieee%3A42020%3Aed-1%3Av1%3Aen">ISO/IEC/IEEE 42020:2019</a></p>
<p>That purpose-centered definition is appropriate for engineered systems. It should not be projected onto every natural system. A climate system and a river system display organized behavior without needing intentions of their own.</p>
<h2 id="a-system-is-not-an-inventory">A system is not an inventory</h2>
<p>Suppose an online publishing operation contains an author, articles, Markdown files, a Git repository, a static-site generator, a deployment service, a domain, search engines, and readers.</p>
<p>The list tells us what might be present. It does not yet explain the system. The explanation begins with relations:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">author ──writes──&gt; article
</span></span><span class="line"><span class="cl">article ──stored as──&gt; Markdown
</span></span><span class="line"><span class="cl">Git ──records──&gt; revision
</span></span><span class="line"><span class="cl">site generator ──transforms──&gt; web page
</span></span><span class="line"><span class="cl">deployment service ──publishes──&gt; site
</span></span><span class="line"><span class="cl">reader ──visits──&gt; page
</span></span><span class="line"><span class="cl">search engine ──indexes──&gt; content
</span></span></code></pre></div><p>The system exists as an organized pattern of dependencies, transformations, permissions, and flows. If those relations disappear, the same objects may remain, but the publishing capability does not.</p>
<p>This is why a system is not simply the sum of its components. Organization is causally relevant.</p>
<h2 id="how-is-a-system-different-from-a-collection">How is a system different from a collection?</h2>
<p>A collection is defined mainly by membership. A system is defined by interdependence and organization.</p>
<table>
  <thead>
      <tr>
          <th>Collection</th>
          <th>System</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Answers which members are included</td>
          <td>Answers how elements interact</td>
      </tr>
      <tr>
          <td>Members may remain independent</td>
          <td>Changes can propagate among elements</td>
      </tr>
      <tr>
          <td>Removing one member may only change the count</td>
          <td>Changing one part may alter overall behavior</td>
      </tr>
      <tr>
          <td>Does not require a shared capacity</td>
          <td>Organization may create a capacity of the whole</td>
      </tr>
  </tbody>
</table>
<p>One hundred chairs in a warehouse form a collection. Players, coaches, rules, training, communication, and matches may form a team system. A contact list is a collection of records. Communication, roles, authority, and feedback can turn a group of people into an operating organization.</p>
<h2 id="boundaries-make-system-analysis-possible">Boundaries make system analysis possible</h2>
<p>Everything is connected to something else. If every remote influence must be included, the system expands until it becomes indistinguishable from the world. A useful analysis therefore defines a boundary.</p>
<p>A boundary answers:</p>
<blockquote>
<p>Which elements and relations belong to the system under study, and which belong to its environment?</p>
</blockquote>
<p>Consider a coffee shop. An analysis of waiting time may include customers, ordering, baristas, equipment, and queue rules. An analysis of profitability may add rent, suppliers, delivery platforms, and pricing. A food-safety analysis may include storage, temperature control, cleaning, and regulation.</p>
<p>The coffee shop has not become three different objects. The analytical boundary changes because the question changes.</p>
<p>INCOSE describes the environment as the part of the outside world that significantly interacts with and affects a system, often as the source of inputs and destination of outputs. Defining the boundary and environment is therefore a basic step in systems thinking.<a href="https://www.incose.org/wp-content/uploads/2026/01/INCOSEContent-410.pdf">INCOSE: Systems Thinking 101</a></p>
<p>The boundary is selected, but it is not arbitrary. A model cannot exclude a major causal influence merely because including it would make the diagram untidy.</p>
<h2 id="open-systems-exchange-with-an-environment">Open systems exchange with an environment</h2>
<p>An environment may provide information, energy, materials, users, prices, legal constraints, threats, and disturbances. A system may return products, decisions, services, waste, risk, and social effects.</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">environment
</span></span><span class="line"><span class="cl">    ↓ input
</span></span><span class="line"><span class="cl">system state and operation
</span></span><span class="line"><span class="cl">    ↓ output
</span></span><span class="line"><span class="cl">changed environment
</span></span><span class="line"><span class="cl">    ↓ new input
</span></span><span class="line"><span class="cl">continued operation
</span></span></code></pre></div><p>Most systems studied in biology, society, organizations, and computing are open in this sense. They depend on continued exchange.</p>
<p>A closed system is often an analytical idealization. It means that particular exchanges can be ignored for a particular purpose, not that the object has no relation to anything outside it.</p>
<h2 id="systems-have-state-and-history">Systems have state and history</h2>
<p>A system is not only an architecture diagram. It occupies states, changes state, and carries effects from its history.</p>
<p>A website may be healthy, building, partially deployed, unavailable because of DNS, or updated in a repository while production still serves an earlier release. The same components and nominal connections can therefore produce different outcomes at different moments.</p>
<p>A simple dynamic description is:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">current state + input + transition rules
</span></span><span class="line"><span class="cl">                    ↓
</span></span><span class="line"><span class="cl">             next state + output
</span></span></code></pre></div><p>Time delays matter. A page may be live before a search engine indexes it. A product may improve before public expectations change. A policy may appear effective in the short term while accumulating a long-term cost.</p>
<p>Ignoring delay turns a dynamic system into a misleading snapshot.</p>
<h2 id="feedback-changes-subsequent-behavior">Feedback changes subsequent behavior</h2>
<p>Feedback occurs when an output or consequence returns as an influence on later system behavior.</p>
<p><strong>Negative feedback</strong> counteracts a deviation. A thermostat detects a falling temperature, turns on heating, and later turns it off as the room returns to range. Negative feedback often supports stability.</p>
<p><strong>Positive feedback</strong> reinforces a change. More engagement may produce more recommendation exposure, which produces more engagement. Positive feedback may create growth, lock-in, polarization, or collapse.</p>
<p>The words <em>positive</em> and <em>negative</em> do not mean good and bad. They describe whether the loop amplifies or counteracts change.</p>
<p>Delayed feedback can cause overshoot and oscillation. Missing feedback can allow a system to optimize a proxy long after the proxy has stopped representing the intended result.</p>
<h2 id="emergence-comes-from-organization">Emergence comes from organization</h2>
<p>“The whole is greater than the sum of its parts” is a familiar systems slogan. It is suggestive but imprecise.</p>
<p>What the inventory omits is organization: spatial arrangement, causal interaction, timing, feedback, and constraints. Those relations allow a whole to display properties that isolated parts do not display.</p>
<ul>
<li>A single vehicle does not constitute a traffic jam; many mutually constraining vehicles can.</li>
<li>A single market participant does not determine a market price; structured interaction among many participants may produce one.</li>
<li>A single neuron does not perform the full cognitive work of a nervous system; organized neural activity supports higher-level capacities.</li>
<li>A molecule does not have the thermodynamic temperature of a macroscopic body; temperature characterizes a collective state.</li>
</ul>
<p>Systems biology likewise studies components in the context of their interactions and the constraints imposed by the whole.<a href="https://plato.stanford.edu/entries/systems-synthetic-biology/">Stanford Encyclopedia of Philosophy: Philosophy of Systems and Synthetic Biology</a></p>
<p>Emergence should not be used as a label for mystery. An emergent property still calls for an explanation of arrangement, interaction, scale, constraint, and time.</p>
<h2 id="does-every-system-have-a-goal">Does every system have a goal?</h2>
<p>No. Purpose, function, and intention must be distinguished.</p>
<p>An engineered payment system is built for stated purposes. An organization may pursue several partly conflicting objectives. A heart performs a biological function that can be explained through physiology and evolution without attributing intention to the organ. A solar system displays lawful behavior without pursuing a goal.</p>
<p>Even in designed systems, four things can diverge:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">stated purpose
</span></span><span class="line"><span class="cl">designer intention
</span></span><span class="line"><span class="cl">metric actually optimized
</span></span><span class="line"><span class="cl">outcome actually produced
</span></span></code></pre></div><p>A recommendation service may claim to improve user satisfaction while optimizing clicks. Higher click-through rates may coexist with lower trust, worse information quality, or compulsive use.</p>
<p>A system should therefore be evaluated by its operation and effects, not only by its declared purpose.</p>
<h2 id="system-structure-mechanism-process-and-model">System, structure, mechanism, process, and model</h2>
<p>These terms answer different questions.</p>
<table>
  <thead>
      <tr>
          <th>Term</th>
          <th>Central question</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>System</td>
          <td>Which elements interact within which boundary, producing what behavior?</td>
      </tr>
      <tr>
          <td>Structure</td>
          <td>How are elements arranged, connected, and layered?</td>
      </tr>
      <tr>
          <td>Mechanism</td>
          <td>Through which causal organization is an outcome produced?</td>
      </tr>
      <tr>
          <td>Process</td>
          <td>In what temporal sequence do activities and transformations occur?</td>
      </tr>
      <tr>
          <td>Function</td>
          <td>What capacity or contribution does a whole or part provide?</td>
      </tr>
      <tr>
          <td>Organization</td>
          <td>How are people, roles, rules, and authority coordinated?</td>
      </tr>
      <tr>
          <td>Network</td>
          <td>What topology is formed by nodes and links?</td>
      </tr>
      <tr>
          <td>Model</td>
          <td>How is the object represented for understanding, explanation, or prediction?</td>
      </tr>
  </tbody>
</table>
<p>A publishing system has a structure of repositories, builders, and servers; a mechanism that converts Markdown into HTML; a process of drafting, review, commit, build, and deployment; and a function of making content reliably accessible.</p>
<p>The system is the object under study. A system model is a selective representation of that object. A diagram may omit details usefully, but success in the diagram does not guarantee success in the world.</p>
<h2 id="systems-entities-and-ontology">Systems, entities, and ontology</h2>
<p>An entity analysis asks which particular object is being referred to. A system analysis asks how entities and relations are organized into behavior over time. An ontology specifies which kinds of entities and relations a domain recognizes.</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">entity: an identifiable object
</span></span><span class="line"><span class="cl">relation: a way objects are connected
</span></span><span class="line"><span class="cl">system: organized objects and relations operating over time
</span></span><span class="line"><span class="cl">ontology: an account of the recognized kinds and relations
</span></span></code></pre></div><p>A system can itself be treated as an entity. A payment system may have an owner, version, status, service boundary, and lifecycle. At a lower level, the same system contains account, order, fraud-control, channel, and settlement subsystems.</p>
<p>Something can therefore be a system at one level and a component of a larger system at another. The relevant level depends on the question.</p>
<p>Related discussions appear in <a href="/en/notes/what-is-an-entity/">“What Is an Entity? Identity, Reference, and AI Systems”</a> and <a href="/en/notes/what-is-ontology/">“What Is Ontology? From What Exists to What AI Can Represent”</a>.</p>
<h2 id="subsystems-and-systems-of-systems">Subsystems and systems of systems</h2>
<p>Complex systems are often nested:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">component
</span></span><span class="line"><span class="cl">    ↓
</span></span><span class="line"><span class="cl">subsystem
</span></span><span class="line"><span class="cl">    ↓
</span></span><span class="line"><span class="cl">system
</span></span><span class="line"><span class="cl">    ↓
</span></span><span class="line"><span class="cl">larger system
</span></span></code></pre></div><p>A payment interface may belong to an order-and-payment subsystem, which belongs to an e-commerce platform, which participates in a broader commercial and logistics environment.</p>
<p>A <strong>system of systems</strong> joins systems that can still operate and be managed with substantial independence. Urban mobility, for example, may involve roads, buses, rail, navigation, ticketing, traffic control, taxis, and ride-hailing platforms. No single component completely determines the whole, yet their interactions shape citywide movement.</p>
<p>This creates coordination problems that cannot be solved by optimizing one subsystem alone.</p>
<h2 id="are-systems-discovered-or-constructed">Are systems discovered or constructed?</h2>
<p>Both descriptions capture part of the truth.</p>
<p>Real interactions, dependencies, feedback loops, and constraints are not invented merely by drawing a boundary. Traffic congestion, ecological exchange, institutional authority, and software calls have real effects.</p>
<p>Yet the choice of boundary, scale, state variables, and level of abstraction depends on what the investigator needs to explain. The same person can participate in biological, family, legal, organizational, economic, and information systems.</p>
<p>A useful position is:</p>
<blockquote>
<p><strong>Interactions are constrained by reality; system boundaries and levels are selected for inquiry and action.</strong></p>
</blockquote>
<p>The model must remain answerable to observed behavior. A convenient boundary that excludes decisive effects is a bad boundary.</p>
<h2 id="system-boundaries-also-carry-values-and-power">System boundaries also carry values and power</h2>
<p>Boundary choices determine what becomes visible.</p>
<p>When a system defines who counts as a user, which outcomes count as benefits, which harms count as externalities, which metric receives optimization pressure, and who may change the rules, it embeds practical and political judgments.</p>
<p>A delivery platform that measures only completed orders per hour may improve its internal efficiency while moving safety risk, waiting pressure, and road danger onto workers and the public.</p>
<p>The system did not eliminate the cost. Its measurement boundary excluded the cost.</p>
<p>Systems analysis therefore asks more than whether an operation is efficient:</p>
<ul>
<li>Efficient for whom?</li>
<li>Which outcomes are measured?</li>
<li>Who absorbs failure and delay?</li>
<li>Who has authority to change the objective or boundary?</li>
<li>Which effects appear only in a larger system?</li>
</ul>
<h2 id="what-is-an-ai-system">What is an AI system?</h2>
<p>An AI system is not synonymous with an AI model.</p>
<p>A deployed language-model application may include:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">users and operators
</span></span><span class="line"><span class="cl">+ interface
</span></span><span class="line"><span class="cl">+ prompts and context
</span></span><span class="line"><span class="cl">+ one or more models
</span></span><span class="line"><span class="cl">+ retrieval and knowledge sources
</span></span><span class="line"><span class="cl">+ memory and state
</span></span><span class="line"><span class="cl">+ tools and external services
</span></span><span class="line"><span class="cl">+ orchestration and control flow
</span></span><span class="line"><span class="cl">+ identity and permissions
</span></span><span class="line"><span class="cl">+ runtime infrastructure
</span></span><span class="line"><span class="cl">+ logging and evaluation
</span></span><span class="line"><span class="cl">+ human review
</span></span><span class="line"><span class="cl">+ outcome feedback
</span></span></code></pre></div><p>The model performs part of the inference. The surrounding system decides what reaches the model, which external state is available, which tools may run, whose authority applies, how outputs are checked, and what changes in the world.</p>
<p>The OECD definition describes an AI system as a machine-based system that infers from inputs how to produce predictions, content, recommendations, or decisions that can influence physical or virtual environments. Its explanatory material describes a model as a core component of such a system.<a href="https://oecd.ai/en/wonk/ai-system-definition-update">OECD: Updated Definition of an AI System</a></p>
<p>NIST emphasizes that AI systems are sociotechnical: their benefits and risks emerge from technical components together with operators, users, other systems, and the social context of deployment.<a href="https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf">NIST AI Risk Management Framework</a></p>
<h2 id="a-model-can-answer-while-the-system-still-fails">A model can answer while the system still fails</h2>
<p>A language model may generate a correct article draft. Publishing that article reliably requires a larger system to:</p>
<ul>
<li>resolve the intended article and language versions;</li>
<li>read repository rules;</li>
<li>modify the correct files;</li>
<li>preserve identity and authorization boundaries;</li>
<li>validate the build;</li>
<li>commit and push with the correct repository identity;</li>
<li>wait for deployment;</li>
<li>verify the live pages, canonical URLs, language links, and sitemap.</li>
</ul>
<p>Failure can occur at any boundary. The model may be correct while retrieval supplies stale facts. The plan may be correct while a tool targets the wrong entity. The code may build locally while deployment fails. The page may be live while search metadata is missing.</p>
<p>Therefore:</p>
<blockquote>
<p><strong>Model capability does not by itself establish system reliability.</strong></p>
</blockquote>
<p>Evaluation must match the system boundary. A model benchmark measures a model under specified conditions. It does not automatically measure the complete product, workflow, organization, or real-world outcome.</p>
<h2 id="a-practical-method-for-analyzing-a-system">A practical method for analyzing a system</h2>
<ol>
<li><strong>State the question.</strong> Are you trying to explain, design, improve, control, or evaluate the system?</li>
<li><strong>Draw the boundary.</strong> What belongs inside, what belongs to the environment, and why?</li>
<li><strong>Identify elements.</strong> Include people, software, physical objects, rules, institutions, data, and resources.</li>
<li><strong>Map relations and flows.</strong> Track information, material, money, authority, and dependency.</li>
<li><strong>Describe states and transitions.</strong> What states can occur, and what changes one state into another?</li>
<li><strong>Find feedback loops.</strong> Which consequences reinforce or counteract earlier behavior?</li>
<li><strong>Examine delays.</strong> When do outputs and consequences become observable?</li>
<li><strong>Separate goals from metrics.</strong> What outcome is intended, and what variable actually receives optimization pressure?</li>
<li><strong>Locate constraints and authority.</strong> Who can change which element, rule, or boundary?</li>
<li><strong>Inspect excluded effects.</strong> Which people, costs, and risks appear only when the boundary expands?</li>
</ol>
<h2 id="a-final-definition">A final definition</h2>
<p>A system can be defined as:</p>
<blockquote>
<p><strong>A system is a bounded whole in which interdependent elements operate through some organization and rules, change state over time, interact with an environment, and thereby produce behavior or capacities at the level of the whole.</strong></p>
</blockquote>
<p>For designed systems, that account must also include purpose, constraints, metrics, authority, and responsibility.</p>
<p>A system is not merely many things placed together, and it is not merely an architecture diagram. It exists as organized interaction: changes in one part affect others, relations alter outcomes, and the condition of the whole constrains what its parts can do.</p>
<p>Identifying entities tells us what is present. Understanding a system tells us how those entities operate together and why the observed result occurs.</p>
<h2 id="references">References</h2>
<ul>
<li><a href="https://www.etymonline.com/word/system">Etymonline: system</a></li>
<li><a href="https://www.iso.org/obp/ui?_escaped_fragment_=iso%3Astd%3Aiso-iec-ieee%3A42020%3Aed-1%3Av1%3Aen">ISO/IEC/IEEE 42020:2019</a></li>
<li><a href="https://www.incose.org/wp-content/uploads/2026/01/INCOSEContent-410.pdf">INCOSE: Systems Thinking 101</a></li>
<li><a href="https://plato.stanford.edu/entries/systems-synthetic-biology/">Stanford Encyclopedia of Philosophy: Philosophy of Systems and Synthetic Biology</a></li>
<li><a href="https://oecd.ai/en/wonk/ai-system-definition-update">OECD: Updated Definition of an AI System</a></li>
<li><a href="https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf">NIST AI Risk Management Framework</a></li>
</ul>
]]></content:encoded></item><item><title>Entities: Identity, Reference, and AI Systems</title><link>https://moonment.net/en/notes/what-is-an-entity/</link><pubDate>Thu, 24 Sep 2026 16:31:31 +0800</pubDate><dc:creator>Moon</dc:creator><guid>https://moonment.net/en/notes/what-is-an-entity/</guid><description>An entity is something treated as having a distinguishable identity. This essay separates mentions, names, classes, records, and real-world referents across philosophy and AI.</description><content:encoded><![CDATA[<p>People, companies, cities, products, documents, events, and fictional characters are very different kinds of things. Yet an information system may treat all of them as entities.</p>
<p>The word is broad because it does not describe a particular material composition. It describes a role within thought, language, or a model:</p>
<blockquote>
<p><strong>An entity is something treated as having a distinguishable identity, so that it can be referred to, described, related to other things, tracked over time, or acted upon.</strong></p>
</blockquote>
<p>An entity need not be tangible. It need not exist independently. It need not even exist in the actual world. What it needs, within a particular domain, is enough identity for the system to treat it as the same subject across multiple statements or operations.</p>
<h2 id="where-does-the-word-entity-come-from">Where does the word “entity” come from?</h2>
<p>English <code>entity</code> comes through Medieval Latin <code>entitas</code>, formed from <code>ens</code>, a being or something that is, which in turn is related to the Latin verb <code>esse</code>, to be. The word therefore carries an ontological background: it concerns something considered as a being or item of existence.<a href="https://www.ahdictionary.com/word/search.html?q=entity">American Heritage Dictionary: entity</a></p>
<p>That history explains why <code>entity</code> can be used so widely. It may refer to a physical object, a person, an institution, an event, a number, a proposition, or another item admitted by a theory. In philosophy, <code>thing</code>, <code>being</code>, <code>entity</code>, and <code>object</code> may all compete for the role of a maximally general term for whatever a system acknowledges.<a href="https://plato.stanford.edu/entries/object/">Stanford Encyclopedia of Philosophy: Object</a></p>
<p>No single list of entities is philosophically neutral. A physicalist ontology, a mathematical ontology, a legal ontology, and a fictional world may recognize different kinds of things. Calling something an entity is therefore both a semantic move and, in many contexts, an ontological commitment.</p>
<h2 id="an-entity-is-not-merely-a-physical-object">An entity is not merely a physical object</h2>
<p>A physical object usually has material structure and some spatial boundary. An entity need not.</p>
<p>A corporation has no single body. It depends on law, records, roles, property, contracts, and continued institutional recognition. A meeting is not a durable object, but it can have participants, a time, a location, an agenda, and an outcome. An account exists only within a platform and its rules, yet the platform must still distinguish one account from another.</p>
<p>All three can function as entities because each can be identified and become the subject of further claims.</p>
<p>The English words <code>entity</code> and <code>substance</code> should also be separated. An entity is any item treated as a being or object of reference. A substance, in major philosophical traditions, is more specifically something taken to exist relatively independently or to bear properties. Events, relations, numbers, and properties may count as entities without counting as substances in that stronger sense.<a href="https://plato.stanford.edu/entries/substance/">Stanford Encyclopedia of Philosophy: Substance</a></p>
<h2 id="reference-does-not-establish-real-existence">Reference does not establish real existence</h2>
<p>Sherlock Holmes can be named, described, compared with other characters, and placed in a network of fictional relations. He is an entity in literary discourse and may be an entity in a knowledge base. None of this makes him a historical person.</p>
<p>At least three questions must therefore remain separate:</p>
<table>
  <thead>
      <tr>
          <th>Level</th>
          <th>Question</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Discourse entity</td>
          <td>Has language introduced something that can be referred to again?</td>
      </tr>
      <tr>
          <td>Model entity</td>
          <td>Does an information system represent it as a distinct item?</td>
      </tr>
      <tr>
          <td>Real-world entity</td>
          <td>Is there sufficient evidence for a corresponding thing in the actual world?</td>
      </tr>
  </tbody>
</table>
<p>An entity can exist at the first two levels without satisfying the third. Plans, hypothetical products, possible events, mistaken identities, and fictional characters all demonstrate why representation and reality must not be collapsed.</p>
<h2 id="identity-is-the-central-problem">Identity is the central problem</h2>
<p>Finding a noun is easy. Determining what makes something the same entity is harder.</p>
<p>A person may change names and addresses while remaining the same person. A corporation may replace every employee and continue as the same legal organization. An article may be revised many times while retaining one publication history. A product may keep its commercial name while its capabilities, components, or terms change substantially.</p>
<p>Different types of entities require different <strong>identity criteria</strong>:</p>
<table>
  <thead>
      <tr>
          <th>Entity type</th>
          <th>Possible basis of identity</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Person</td>
          <td>legal records, bodily and biographical continuity</td>
      </tr>
      <tr>
          <td>Corporation</td>
          <td>legal registration, organizational and contractual continuity</td>
      </tr>
      <tr>
          <td>Document</td>
          <td>identifier, provenance, content hash, or version history</td>
      </tr>
      <tr>
          <td>Article</td>
          <td>authorship and publication lineage, slug, revision history</td>
      </tr>
      <tr>
          <td>Product model</td>
          <td>model definition, capability boundary, specification</td>
      </tr>
      <tr>
          <td>Commercial item</td>
          <td>SKU, serial number, batch, or unit of sale</td>
      </tr>
      <tr>
          <td>Online account</td>
          <td>platform, stable account ID, control and authentication</td>
      </tr>
      <tr>
          <td>Event</td>
          <td>participants, time, place, and occurrence structure</td>
      </tr>
  </tbody>
</table>
<p>A name is not an identity. Two people can share a name, and one person can use several names. A set of properties is not automatically an identity either. Properties change, records conflict, and two objects may resemble one another closely without being the same object.</p>
<p>A useful entity representation usually needs more than a label:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">stable identifier
</span></span><span class="line"><span class="cl">+ entity type
</span></span><span class="line"><span class="cl">+ attributes
</span></span><span class="line"><span class="cl">+ relationships
</span></span><span class="line"><span class="cl">+ temporal state
</span></span><span class="line"><span class="cl">+ provenance
</span></span><span class="line"><span class="cl">+ confidence in the identity match
</span></span></code></pre></div><h2 id="how-does-ontology-relate-to-entities">How does ontology relate to entities?</h2>
<p>An entity does not become a useful modeling unit in isolation. A system must decide what kinds of entities it recognizes, which attributes they may have, which relations may connect them, and what counts as persistence or change.</p>
<p>Those decisions form part of an ontology.</p>
<blockquote>
<p><strong>An entity is a particular item recognized within a domain; an ontology states which kinds of items the domain recognizes and how they may be organized.</strong></p>
</blockquote>
<p>“Shanghai” may be stored as a string in a simple customer table. In a geographical knowledge system, it may be represented as an entity with coordinates, administrative status, districts, and relationships to other places. The difference depends on whether the system needs to identify Shanghai independently and reason about it.</p>
<p>Philosophical ontology asks what exists and how different kinds of beings exist. Computational ontology turns a domain commitment into an explicit model of classes, entities, properties, relations, and constraints. A fuller account appears in <a href="/en/notes/what-is-ontology/">“What Is Ontology? From What Exists to Models AI Can Use”</a>. Here ontology matters because it supplies the type system and identity conditions within which entities can be distinguished.</p>
<h2 id="what-does-entity-mean-in-ai">What does “entity” mean in AI?</h2>
<p>In AI, an entity is usually an object distinguished from text, images, records, or an environment because the system needs to understand, retrieve, remember, reason about, or act on it.</p>
<p>The term changes meaning across tasks. Named-entity recognition, entity linking, entity resolution, knowledge graphs, computer vision, databases, and AI agents do not operate at exactly the same level.</p>
<h2 id="named-entity-recognition-finds-mentions-not-verified-objects">Named-entity recognition finds mentions, not verified objects</h2>
<p>Consider the sentence:</p>
<blockquote>
<p>Apple plans to announce a new phone in Shanghai tomorrow.</p>
</blockquote>
<p>A named-entity recognition system may label:</p>
<ul>
<li><code>Apple</code> as an organization;</li>
<li><code>Shanghai</code> as a location;</li>
<li><code>tomorrow</code> as a date.</li>
</ul>
<p>At this stage, the system has identified <strong>mentions</strong>: spans of language that appear to refer to named or otherwise categorized entities. In engineering terms, a conventional NER component predicts labeled token spans. spaCy, for example, describes its entity recognizer as identifying non-overlapping labeled spans.<a href="https://spacy.io/api/entityrecognizer/">spaCy: EntityRecognizer</a></p>
<p>The output does not yet prove that <code>Apple</code> refers to Apple Inc. rather than a different organization, a title, or an annotation mistake. Nor does it establish that the announced phone is a particular product with a known identity.</p>
<p>The distinction is fundamental:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">entity mention ≠ entity identity
</span></span><span class="line"><span class="cl">entity name ≠ unique referent
</span></span><span class="line"><span class="cl">entity type ≠ proof of existence
</span></span></code></pre></div><h2 id="entity-linking-connects-a-mention-to-a-canonical-entity">Entity linking connects a mention to a canonical entity</h2>
<p>Entity linking attempts to determine which known entity a mention refers to.</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">“Apple” in a sentence
</span></span><span class="line"><span class="cl">        ↓ disambiguation
</span></span><span class="line"><span class="cl">Apple Inc.
</span></span><span class="line"><span class="cl">        ↓ normalization
</span></span><span class="line"><span class="cl">knowledge-base ID: company/apple-inc
</span></span></code></pre></div><p>This requires at least two forms of reasoning:</p>
<ul>
<li><strong>Disambiguation:</strong> Which entity with this name fits the context?</li>
<li><strong>Coreference and normalization:</strong> Do <code>Apple</code>, <code>Apple Inc.</code>, and <code>the Cupertino company</code> refer to the same entity here?</li>
</ul>
<p>Linking converts a piece of language into an addressable object in a knowledge system. It is still fallible. The selected knowledge-base entry may be wrong, duplicated, incomplete, or out of date.</p>
<h2 id="entity-resolution-asks-whether-records-describe-the-same-thing">Entity resolution asks whether records describe the same thing</h2>
<p>Entity linking usually begins with language and a knowledge base. <strong>Entity resolution</strong> often begins with records.</p>
<p>A customer system may contain:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">Li Ming, Shanghai, 138****1234
</span></span><span class="line"><span class="cl">Ming Li, 上海市, 138****1234
</span></span><span class="line"><span class="cl">李明, Pudong, 138****1234
</span></span></code></pre></div><p>Do these records describe one person, two people, or three? Shared fields provide evidence, but they do not make the answer automatic. Phone numbers can be reassigned, addresses can be shared, and names can collide.</p>
<p>Entity resolution may involve:</p>
<ul>
<li>deduplicating records;</li>
<li>matching aliases and transliterations;</li>
<li>detecting that one entity has split or merged in a source system;</li>
<li>preserving conflicting claims rather than forcing a premature merge;</li>
<li>recording why two records were considered the same.</li>
</ul>
<p>This is an identity decision under uncertainty. A useful system keeps the evidence and confidence behind the match instead of treating similarity as certainty.</p>
<h2 id="knowledge-graphs-place-entities-in-a-network-of-claims">Knowledge graphs place entities in a network of claims</h2>
<p>In a knowledge graph, entities are commonly represented as nodes connected by typed relations:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">Apple Inc. ──headquartered in──&gt; Cupertino
</span></span><span class="line"><span class="cl">Apple Inc. ──released──&gt; iPhone
</span></span><span class="line"><span class="cl">iPhone ──instance of──&gt; smartphone product
</span></span></code></pre></div><p>This representation distinguishes:</p>
<ul>
<li>entities such as Apple Inc., Cupertino, and an iPhone model;</li>
<li>classes such as company, city, and product;</li>
<li>attributes such as dates and names;</li>
<li>relations such as <code>headquartered in</code> and <code>released</code>.</li>
</ul>
<p>RDF represents claims as subject–predicate–object triples. Its notion of a resource is deliberately broad: a resource may denote a physical thing, a document, an abstract concept, or another item in the universe of discourse.<a href="https://www.w3.org/TR/rdf11-concepts/">W3C: RDF 1.1 Concepts and Abstract Syntax</a></p>
<p>An ontology supplies general rules, while a knowledge graph contains particular claims:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">ontology: a company may release a product
</span></span><span class="line"><span class="cl">claim: Apple Inc. released a particular iPhone model
</span></span></code></pre></div><p>A graph node is not automatically a verified real-world object. It may represent a class, a fictional entity, a planned object, an uncertain hypothesis, or an erroneous record. Provenance and status remain necessary.</p>
<h2 id="concepts-classes-entities-identifiers-and-records">Concepts, classes, entities, identifiers, and records</h2>
<p>Several layers are easily confused:</p>
<table>
  <thead>
      <tr>
          <th>Layer</th>
          <th>What does it provide?</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Word or phrase</td>
          <td>an expression used in language</td>
      </tr>
      <tr>
          <td>Name</td>
          <td>a conventional way to refer to something</td>
      </tr>
      <tr>
          <td>Concept</td>
          <td>a general structure used to understand things</td>
      </tr>
      <tr>
          <td>Class</td>
          <td>a modeled category of possible members</td>
      </tr>
      <tr>
          <td>Entity</td>
          <td>the particular item currently referred to</td>
      </tr>
      <tr>
          <td>Identifier</td>
          <td>a stable handle used by a system</td>
      </tr>
      <tr>
          <td>Record</td>
          <td>stored claims about an entity</td>
      </tr>
      <tr>
          <td>Real-world referent</td>
          <td>whatever, if anything, exists beyond the model</td>
      </tr>
  </tbody>
</table>
<p><code>Company</code> may be a concept and a class. Apple Inc. may be an entity. <code>Apple</code> may be a name or mention. <code>company_001</code> may be an internal identifier. A row in a database may contain claims about the company. None of these layers is identical to the organization itself.</p>
<p>One entity may have many names and records. One record may accidentally combine facts about several entities. A unique database key guarantees uniqueness inside a table; it does not prove that the row corresponds correctly to one real-world thing.</p>
<h2 id="when-should-something-be-modeled-as-an-entity">When should something be modeled as an entity?</h2>
<p>Not every noun phrase needs its own entity. Six questions help:</p>
<ol>
<li><strong>Does it need a distinct identity?</strong> Must the system distinguish this item from similar items?</li>
<li><strong>Will it be referred to repeatedly?</strong> Will multiple documents, records, or tasks mention it?</li>
<li><strong>Does it have its own attributes?</strong> Must the system store a status, date, location, or owner?</li>
<li><strong>Does it participate in relationships?</strong> Must it be connected to other objects?</li>
<li><strong>Does it persist through time?</strong> Must the system recognize it after some properties change?</li>
<li><strong>Can the system act on it?</strong> Will it be queried, edited, sent, authorized, purchased, deleted, or monitored?</li>
</ol>
<p>If most answers are yes, an entity is often appropriate. Otherwise a literal value, label, or temporary span may be enough.</p>
<p>Entity modeling is not free. Every entity type introduces identity rules, lifecycle questions, merge and split behavior, provenance requirements, and access-control consequences.</p>
<h2 id="can-events-states-and-intentions-become-entities">Can events, states, and intentions become entities?</h2>
<p>Grammatical categories do not map directly onto model categories. Nouns do not always denote entities, and verbs do not always remain mere relations.</p>
<p>“Maya signed Contract C36” can be represented as a simple relation:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">Maya ──signed──&gt; Contract C36
</span></span></code></pre></div><p>If the system must record the date, location, version, witnesses, method, and legal status, the signing can become an event entity:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">Signing Event E1024
</span></span><span class="line"><span class="cl">├── signer: Maya
</span></span><span class="line"><span class="cl">├── object: Contract C36
</span></span><span class="line"><span class="cl">├── time: 2026-09-24
</span></span><span class="line"><span class="cl">├── place: Shanghai
</span></span><span class="line"><span class="cl">└── status: effective
</span></span></code></pre></div><p>An intention can likewise be modeled as a mental-state entity when the system needs to record whose intention it is, what outcome it concerns, when it was inferred, which evidence supports the inference, and how uncertain the interpretation remains.</p>
<p>Reification makes an event, relation, or state available for further description. It is useful when the system needs that detail, but it should not be mistaken for a discovery that the world naturally divides itself in exactly that form.</p>
<h2 id="how-do-large-language-models-handle-entities">How do large language models handle entities?</h2>
<p>A standard large language model receives tokens and computes distributed representations shaped by training data and current context. It can learn that <code>Apple</code> often occurs near company, product, iPhone, and Cupertino, and it can frequently disambiguate the word from the fruit.</p>
<p>This competence does not imply that the model contains a single, explicit, canonical Apple Inc. record comparable to a carefully maintained knowledge base. Entity information may be distributed across parameters and reconstructed probabilistically in context.</p>
<p>Consequently, a language model may:</p>
<ul>
<li>resolve an entity correctly in one context and confuse it in another;</li>
<li>conflate a company, its brand, its products, and its website;</li>
<li>recall an old property without knowing that it has changed;</li>
<li>invent a plausible person, paper, organization, or product;</li>
<li>answer without a stable source for the entity claim.</li>
</ul>
<p>For tasks that require reliable action, model output is usually only one layer:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">language model
</span></span><span class="line"><span class="cl">+ domain ontology
</span></span><span class="line"><span class="cl">+ entity IDs and resolution
</span></span><span class="line"><span class="cl">+ current state and provenance
</span></span><span class="line"><span class="cl">+ time-aware records
</span></span><span class="line"><span class="cl">+ permissions and action rules
</span></span></code></pre></div><p>The language model interprets open-ended language. The ontology defines possible types and relations. Entity services determine identity. Data sources establish current state. Authorization determines which operations are permitted.</p>
<h2 id="why-do-ai-agents-need-explicit-entity-identity">Why do AI agents need explicit entity identity?</h2>
<p>Consider the instruction:</p>
<blockquote>
<p>Update yesterday’s article about needs on Moonment.</p>
</blockquote>
<p>The request contains several unresolved references:</p>
<table>
  <thead>
      <tr>
          <th>Expression</th>
          <th>Entity question</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>yesterday</td>
          <td>Which timezone and time interval?</td>
      </tr>
      <tr>
          <td>the article about needs</td>
          <td>Which document and slug?</td>
      </tr>
      <tr>
          <td>update</td>
          <td>Edit a draft, commit a repository, or publish a website?</td>
      </tr>
      <tr>
          <td>Moonment</td>
          <td>Which project, repository, deployment, and domain?</td>
      </tr>
      <tr>
          <td>article version</td>
          <td>Chinese, English, or both?</td>
      </tr>
      <tr>
          <td>requesting user</td>
          <td>Which permissions and prior authorization apply?</td>
      </tr>
  </tbody>
</table>
<p>The agent may understand the general intention while still acting on the wrong file, project, account, language version, or deployment.</p>
<p>A safer path is:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">natural-language request
</span></span><span class="line"><span class="cl">        ↓
</span></span><span class="line"><span class="cl">mention and reference detection
</span></span><span class="line"><span class="cl">        ↓
</span></span><span class="line"><span class="cl">type assignment and coreference
</span></span><span class="line"><span class="cl">        ↓
</span></span><span class="line"><span class="cl">entity linking and resolution
</span></span><span class="line"><span class="cl">        ↓
</span></span><span class="line"><span class="cl">relation, event, and intent interpretation
</span></span><span class="line"><span class="cl">        ↓
</span></span><span class="line"><span class="cl">state, provenance, and authorization checks
</span></span><span class="line"><span class="cl">        ↓
</span></span><span class="line"><span class="cl">action on the resolved entity
</span></span></code></pre></div><p>Intent specifies the change the user appears to seek. Entity resolution determines what the intended action applies to. Authorization determines whether the action may be performed. None can substitute for the others.</p>
<h2 id="can-ai-establish-that-an-entity-really-exists">Can AI establish that an entity really exists?</h2>
<p>Language alone cannot establish existence.</p>
<p>An AI system can estimate that a phrase is probably a person’s name, that context probably indicates a company, or that a mention probably links to a known record. Real-world verification requires additional evidence, such as:</p>
<ul>
<li>an authoritative registry or primary source;</li>
<li>a current database record;</li>
<li>a file that actually exists in the relevant filesystem;</li>
<li>a live website or API;</li>
<li>an authenticated identity and permission system;</li>
<li>a sensor observation or human confirmation.</li>
</ul>
<p>The following claims are distinct:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">the text mentions something
</span></span><span class="line"><span class="cl">≠ the model identified it correctly
</span></span><span class="line"><span class="cl">≠ the system linked the correct record
</span></span><span class="line"><span class="cl">≠ the record is accurate and current
</span></span><span class="line"><span class="cl">≠ the corresponding thing exists now
</span></span><span class="line"><span class="cl">≠ the user is authorized to act on it
</span></span></code></pre></div><p>Model confidence is not proof of existence. It describes a model’s judgment under particular inputs and assumptions. It does not replace provenance or verification.</p>
<h2 id="a-checklist-for-entity-design-in-ai-systems">A checklist for entity design in AI systems</h2>
<ol>
<li>Which entity types does the system recognize, and why?</li>
<li>Are mentions, names, classes, entities, identifiers, and records kept separate?</li>
<li>What establishes identity for each entity type?</li>
<li>How are aliases, namesakes, renaming, and duplicate records handled?</li>
<li>Do attributes and relations carry time and provenance?</li>
<li>Are events and changing states being flattened into misleading static fields?</li>
<li>Are extraction, linking, resolution, and real-world verification separate stages?</li>
<li>Can the system distinguish fictional, planned, hypothetical, and actual entities?</li>
<li>What happens to historical claims when entities merge, split, or change type?</li>
<li>Before acting, how does the system verify identity, current state, and permission?</li>
</ol>
<p>An entity-rich system can still be unreliable. Without identity criteria, provenance, temporal state, and authorization, it merely attaches confident-looking labels to uncertain referents.</p>
<h2 id="entities-connect-language-to-action">Entities connect language to action</h2>
<p>Entities form an interface between language, knowledge, and operations:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">objects, events, and states in a domain
</span></span><span class="line"><span class="cl">                ↓
</span></span><span class="line"><span class="cl">names and descriptions in language
</span></span><span class="line"><span class="cl">                ↓
</span></span><span class="line"><span class="cl">entity, attribute, and relation recognition
</span></span><span class="line"><span class="cl">                ↓
</span></span><span class="line"><span class="cl">links to records and external evidence
</span></span><span class="line"><span class="cl">                ↓
</span></span><span class="line"><span class="cl">interpretation of needs and intentions
</span></span><span class="line"><span class="cl">                ↓
</span></span><span class="line"><span class="cl">authorized action and recorded outcomes
</span></span></code></pre></div><p>An entity answers, “Which particular thing are we talking about?” An attribute answers, “What is it like?” A relation answers, “How is it connected?” An event answers, “What happened?” An intention answers, “What outcome is an agent trying to bring about?”</p>
<p>The most serious entity error in AI is not failing to label a noun. It is treating an ambiguous expression as though it already named one verified, current, and actionable object.</p>
<h2 id="references">References</h2>
<ul>
<li><a href="https://www.ahdictionary.com/word/search.html?q=entity">American Heritage Dictionary: entity</a></li>
<li><a href="https://plato.stanford.edu/entries/object/">Stanford Encyclopedia of Philosophy: Object</a></li>
<li><a href="https://plato.stanford.edu/entries/substance/">Stanford Encyclopedia of Philosophy: Substance</a></li>
<li><a href="https://plato.stanford.edu/entries/logic-ontology/">Stanford Encyclopedia of Philosophy: Logic and Ontology</a></li>
<li><a href="https://spacy.io/api/entityrecognizer/">spaCy: EntityRecognizer</a></li>
<li><a href="https://spacy.io/usage/linguistic-features#named-entities">spaCy: Linguistic Features—Named Entity Recognition</a></li>
<li><a href="https://www.w3.org/TR/rdf11-concepts/">W3C: RDF 1.1 Concepts and Abstract Syntax</a></li>
<li><a href="https://www.w3.org/TR/owl2-primer/">W3C: OWL 2 Web Ontology Language Primer</a></li>
</ul>
]]></content:encoded></item><item><title>AI Reasoning and Action: From Model Generation to Agent Execution</title><link>https://moonment.net/en/notes/ai-reasoning-and-action/</link><pubDate>Fri, 18 Sep 2026 15:20:00 +0800</pubDate><dc:creator>Moon</dc:creator><guid>https://moonment.net/en/notes/ai-reasoning-and-action/</guid><description>A functional account of language-model reasoning, the limits of visible chains of thought, and the architecture that turns a model into an agent acting through tools.</description><content:encoded><![CDATA[<blockquote>
<p><strong>Series: Thinking, Intention, and Action (2/4).</strong> Previous: <a href="/en/notes/human-thinking/">Human Thinking</a>; next: <a href="/en/notes/ai-user-intent-inference/">User Intent in AI</a></p>
</blockquote>
<h2 id="what-does-it-mean-to-say-that-ai-thinks">What Does It Mean to Say That AI “Thinks”?</h2>
<p>The claim that an AI system thinks can refer to three different questions:</p>
<ol>
<li>Can it perform tasks that require reasoning, planning, comparison, and judgment?</li>
<li>Does its computation contain internal processes that deserve the functional name <em>thinking</em>?</li>
<li>Does it possess consciousness, subjective experience, understanding, or intentions like a person?</li>
</ol>
<p>The first question has an empirical answer: present systems can perform many tasks that previously required human thought.</p>
<p>The second supports a qualified functional definition:</p>
<blockquote>
<p><strong>AI reasoning is the computational transformation of inputs, context, learned parameters, and tool observations into predictions, judgments, plans, and selected outputs.</strong></p>
</blockquote>
<p>The third does not follow from performance. Producing a proof, explaining a concept, or planning a project does not establish that a system experiences its activity or understands it in the way a person does.</p>
<p>The relevant distinctions are:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">behavioral competence
</span></span><span class="line"><span class="cl">≠
</span></span><span class="line"><span class="cl">computational mechanism
</span></span><span class="line"><span class="cl">≠
</span></span><span class="line"><span class="cl">subjective experience
</span></span></code></pre></div><p>This article concerns the first two.</p>
<h2 id="the-base-operation-of-a-language-model">The Base Operation of a Language Model</h2>
<p>A language model is trained to predict a token from the tokens that precede it:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">P(next token | current context, model parameters)
</span></span></code></pre></div><p>Training adjusts a large collection of parameters so that the model becomes sensitive to statistical structure across words, syntax, genres, factual statements, arguments, programs, and patterns of explanation.</p>
<p>At generation time, the simplified cycle is:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">read the current context
</span></span><span class="line"><span class="cl">→ score possible next tokens
</span></span><span class="line"><span class="cl">→ select one token
</span></span><span class="line"><span class="cl">→ append it to the context
</span></span><span class="line"><span class="cl">→ repeat
</span></span></code></pre></div><p>Calling this “next-token prediction” is accurate but incomplete as an explanation of capability. Predicting the continuation of a proof, a program, or a multistep plan can require internal representations that track relations extending far beyond the next word.</p>
<p>The open scientific question is how stable and general those representations are. A model may exhibit a usable concept in one setting, fail after a small reformulation, or rely on a shortcut that worked in the training distribution.</p>
<h2 id="why-prediction-can-produce-reasoning">Why Prediction Can Produce Reasoning</h2>
<p>Human language contains the products of reasoning and many traces of its process: definitions, proofs, disagreements, plans, diagnoses, corrections, and counterexamples. Learning to predict this material exposes a model to recurring structures such as:</p>
<ul>
<li>relevant versus irrelevant evidence;</li>
<li>premises and conclusions;</li>
<li>causes and effects;</li>
<li>goals, constraints, and plans;</li>
<li>programs and execution traces;</li>
<li>claims and objections;</li>
<li>errors and revisions.</li>
</ul>
<p>Large models can consequently perform deduction, induction, analogy, causal explanation, program simulation, and task decomposition to useful degrees.</p>
<p>These abilities remain uneven. Fluent language can hide an invalid inference. Long dependency chains can fail. A familiar template can produce the right answer without a general method, while a slightly unfamiliar case defeats the same model.</p>
<p>It is therefore unsafe to infer reliable reasoning merely from the presence of reasoning-shaped prose.</p>
<h2 id="what-happens-between-prompt-and-output">What Happens Between Prompt and Output?</h2>
<p>Input is divided into tokens and converted into vector representations. A Transformer repeatedly uses attention and nonlinear transformations to construct context-sensitive internal states. The final layers assign scores to possible next tokens.</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">prompt and context
</span></span><span class="line"><span class="cl">→ token and position representations
</span></span><span class="line"><span class="cl">→ attention across relevant positions
</span></span><span class="line"><span class="cl">→ layered internal transformations
</span></span><span class="line"><span class="cl">→ distribution over outputs
</span></span><span class="line"><span class="cl">→ generated continuation
</span></span></code></pre></div><p>Those internal states are not a transcript written in ordinary language. Researchers can probe activations, attention patterns, and latent representations, but there is no simple one-to-one mapping from a single unit to a complete thought.</p>
<p>A model may also generate a step-by-step explanation. Such text can help decompose a problem and make an answer easier to evaluate. It should not be treated as a complete scan of the computation that caused the answer. Experiments have shown that chains of thought can omit influential cues and rationalize a result after the fact. <a href="https://arxiv.org/abs/2305.04388">Turpin et al., <em>Language Models Don&rsquo;t Always Say What They Think</em></a></p>
<blockquote>
<p><strong>A verbal rationale is an interface for work and evaluation, not privileged access to every causal step inside the model.</strong></p>
</blockquote>
<h2 id="from-prediction-to-instruction-following">From Prediction to Instruction Following</h2>
<p>A base model primarily learns what text is likely to follow other text. An assistant must also learn how a request should guide its behavior.</p>
<p>A common development pipeline includes:</p>
<ul>
<li>large-scale pretraining;</li>
<li>supervised examples of instruction following;</li>
<li>optimization from human or model feedback;</li>
<li>runtime system instructions and tool protocols.</li>
</ul>
<p>GPT-3 demonstrated broad in-context task performance from examples and instructions. InstructGPT showed that scale alone does not guarantee alignment with user requests and that instruction tuning plus human feedback can substantially redirect behavior. <a href="https://arxiv.org/abs/2005.14165">GPT-3</a> · <a href="https://arxiv.org/abs/2203.02155">InstructGPT</a></p>
<p>An assistant&rsquo;s response is therefore produced by more than the final user sentence. It depends on learned parameters, system rules, conversation history, visible environment, tool results, and decoding choices.</p>
<h2 id="interpretation-inference-decision-and-action">Interpretation, Inference, Decision, and Action</h2>
<p>Four stages should be kept distinct:</p>
<table>
  <thead>
      <tr>
          <th>Stage</th>
          <th>Governing question</th>
          <th>Typical product</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Interpretation</td>
          <td>What task is being requested?</td>
          <td>Task model, constraints, candidate meanings</td>
      </tr>
      <tr>
          <td>Inference</td>
          <td>What follows from the available evidence?</td>
          <td>Judgments and intermediate conclusions</td>
      </tr>
      <tr>
          <td>Decision</td>
          <td>Which option should be selected?</td>
          <td>Plan, priority, next step</td>
      </tr>
      <tr>
          <td>Action</td>
          <td>How will external state change?</td>
          <td>Tool call, file edit, message, transaction</td>
      </tr>
  </tbody>
</table>
<p>A model can recommend an action without executing it. A system can execute a tool call after weak reasoning. Separating the stages makes failures diagnosable.</p>
<p>“This file appears redundant” is a judgment. “Deleting it will recover space” is a proposed consequence. “Delete it now” is a decision. “The user authorized deletion” is a fact about permission. None of these substitutes for the others.</p>
<h2 id="how-a-model-becomes-an-agent">How a Model Becomes an Agent</h2>
<p>A language model accepts context and emits a continuation. A persistent agent normally requires additional machinery:</p>
<ul>
<li><strong>task state</strong> to record the objective and current progress;</li>
<li><strong>planning</strong> to decompose work into executable steps;</li>
<li><strong>tools</strong> for search, files, code, browsers, and services;</li>
<li><strong>memory</strong> for results, commitments, and stable conventions;</li>
<li><strong>observation</strong> of tool output and environmental state;</li>
<li><strong>permissions</strong> that determine which actions are allowed;</li>
<li><strong>feedback and termination rules</strong> that define completion or revision.</li>
</ul>
<p>The operating loop is:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">observe
</span></span><span class="line"><span class="cl">→ interpret the state
</span></span><span class="line"><span class="cl">→ choose a next action
</span></span><span class="line"><span class="cl">→ invoke a tool
</span></span><span class="line"><span class="cl">→ read the result
</span></span><span class="line"><span class="cl">→ update the plan
</span></span><span class="line"><span class="cl">→ continue or stop
</span></span></code></pre></div><p>ReAct formalized a useful version of this pattern by interleaving reasoning traces with actions and environmental observations. <a href="https://arxiv.org/abs/2210.03629">ReAct</a></p>
<p>The acting unit is therefore not the language model in isolation. It is the assembled system of model, tools, state, permissions, and execution environment.</p>
<h2 id="does-an-ai-agent-have-goals-or-intentions">Does an AI Agent Have Goals or Intentions?</h2>
<p>Engineered systems can contain objective functions, reward signals, task descriptions, and stopping criteria. These should not be collapsed into human desire or practical intention.</p>
<table>
  <thead>
      <tr>
          <th>Level</th>
          <th>Meaning</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Training objective</td>
          <td>Mathematical quantity optimized during training</td>
      </tr>
      <tr>
          <td>System objective</td>
          <td>Task the product or agent is designed to perform</td>
      </tr>
      <tr>
          <td>Current assignment</td>
          <td>Work specified in the present context</td>
      </tr>
      <tr>
          <td>Generated plan</td>
          <td>Proposed subgoals and steps</td>
      </tr>
      <tr>
          <td>Human intention</td>
          <td>A person&rsquo;s purpose, commitment, and orientation toward action</td>
      </tr>
  </tbody>
</table>
<p>When a model writes, “I will inspect the files first,” the sentence can function as a report of the next operation. Its first-person grammar does not establish a private human-like intention.</p>
<p>This is why an agent can display sustained goal-directed behavior while still requiring external authorization, supervision, and an accountable human or institution.</p>
<h2 id="action-requires-feedback">Action Requires Feedback</h2>
<p>A plan produced once cannot guarantee contact with reality. Tools fail, pages change, files disappear, and new evidence defeats earlier assumptions.</p>
<p>Reliable action therefore has a closed loop:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">form a hypothesis
</span></span><span class="line"><span class="cl">→ act
</span></span><span class="line"><span class="cl">→ observe the result
</span></span><span class="line"><span class="cl">→ compare actual and expected state
</span></span><span class="line"><span class="cl">→ revise the interpretation or plan
</span></span><span class="line"><span class="cl">→ act again
</span></span></code></pre></div><p>Without observation, action remains a description inside language. With observation, the system can test whether external state actually changed.</p>
<p>The evidence must also be labeled correctly:</p>
<ul>
<li>generating a command is not executing it;</li>
<li>passing a build is not visual acceptance;</li>
<li>pushing a repository is not proof of deployment;</li>
<li>silence from a user is not authorization for a consequential action.</li>
</ul>
<h2 id="characteristic-failure-modes">Characteristic Failure Modes</h2>
<h3 id="fluency-conceals-weak-evidence">Fluency conceals weak evidence</h3>
<p>A polished answer and a well-supported answer are different achievements.</p>
<h3 id="context-is-partial">Context is partial</h3>
<p>The model can use only the files, messages, tool outputs, and environmental state made available to it. An omitted fact can reverse the correct decision.</p>
<h3 id="long-tasks-lose-state">Long tasks lose state</h3>
<p>Extended work needs checkpoints, external records, and explicit completion criteria. Otherwise a system may repeat steps, omit requirements, or report a plan as a result.</p>
<h3 id="tool-output-still-needs-interpretation">Tool output still needs interpretation</h3>
<p>Search results, webpages, logs, and documents can be incomplete, stale, mistaken, or adversarial. Retrieval changes the evidence set; it does not guarantee truth.</p>
<h3 id="objectives-conflict">Objectives conflict</h3>
<p>User requests, system rules, physical constraints, and local subgoals may point in different directions. Reliable behavior requires detecting and resolving conflict rather than treating every instruction-like string as authoritative.</p>
<h3 id="consequences-belong-to-the-full-system">Consequences belong to the full system</h3>
<p>A model selects a call, an executor changes external state, a platform grants access, and people or institutions assign responsibility. Evaluating only the generated text misses most of the action chain.</p>
<h2 id="how-to-verify-that-an-ai-completed-a-task">How to Verify That an AI Completed a Task</h2>
<p>Confidence should come from evidence at each layer:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">Did it identify the correct object?
</span></span><span class="line"><span class="cl">→ Did it obtain sufficient evidence?
</span></span><span class="line"><span class="cl">→ Does the inference support the conclusion?
</span></span><span class="line"><span class="cl">→ Was the action authorized?
</span></span><span class="line"><span class="cl">→ Did the tool actually run?
</span></span><span class="line"><span class="cl">→ Did external state change as intended?
</span></span><span class="line"><span class="cl">→ Does the outcome satisfy the original objective?
</span></span></code></pre></div><p>For consequential or extended tasks, the system should also preserve recoverable intermediate states so that mistakes can be inspected and reversed.</p>
<h2 id="conclusion">Conclusion</h2>
<p>AI reasoning can be described functionally as the transformation of input, context, and observations into judgments, plans, and selected outputs. This creates real functional comparisons with human thinking, but it does not establish identical consciousness or experience.</p>
<p>AI action belongs to a larger architecture:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">model
</span></span><span class="line"><span class="cl">+ context
</span></span><span class="line"><span class="cl">+ tools
</span></span><span class="line"><span class="cl">+ state
</span></span><span class="line"><span class="cl">+ permissions
</span></span><span class="line"><span class="cl">+ environmental feedback
</span></span></code></pre></div><p>The central questions are therefore not limited to what answer the model generated. They include what evidence it used, how it checked a plan, who authorized execution, what the tools changed, how the result was verified, and how errors update the next cycle.</p>
<blockquote>
<p><strong>AI reasoning computes candidate courses of action. AI agency begins when those computations are connected to authorized tools, observable consequences, and correction through feedback.</strong></p>
</blockquote>
<h2 id="sources">Sources</h2>
<ul>
<li><a href="https://arxiv.org/abs/2005.14165">Brown et al., <em>Language Models are Few-Shot Learners</em></a></li>
<li><a href="https://arxiv.org/abs/2203.02155">Ouyang et al., <em>Training Language Models to Follow Instructions with Human Feedback</em></a></li>
<li><a href="https://arxiv.org/abs/2210.03629">Yao et al., <em>ReAct: Synergizing Reasoning and Acting in Language Models</em></a></li>
<li><a href="https://arxiv.org/abs/2305.04388">Turpin et al., <em>Language Models Don&rsquo;t Always Say What They Think</em></a></li>
</ul>
]]></content:encoded></item><item><title>The Human-AI Action Loop: Interpretation, Coordination, and Feedback</title><link>https://moonment.net/en/notes/human-ai-joint-action-loop/</link><pubDate>Fri, 18 Sep 2026 15:20:00 +0800</pubDate><dc:creator>Moon</dc:creator><guid>https://moonment.net/en/notes/human-ai-joint-action-loop/</guid><description>People express partial intentions, AI systems construct revisable task models, and action plus feedback lets both sides correct goals, plans, and evidence.</description><content:encoded><![CDATA[<blockquote>
<p><strong>Series: Thinking, Intention, and Action (4/4).</strong> Start with <a href="/en/notes/human-thinking/">How Human Thinking Works</a></p>
</blockquote>
<h2 id="collaboration-is-more-than-prompt-and-response">Collaboration Is More Than Prompt and Response</h2>
<p>Human–AI interaction is often pictured as:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">human writes a prompt
</span></span><span class="line"><span class="cl">→ AI returns an answer
</span></span></code></pre></div><p>That picture captures a message exchange. It leaves out why the request arose, how the system selected an interpretation, who authorized an external action, and how the result changes the next decision.</p>
<p>A fuller model is:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">human situation, need, and purpose
</span></span><span class="line"><span class="cl">→ provisional intention
</span></span><span class="line"><span class="cl">→ linguistic request
</span></span><span class="line"><span class="cl">→ AI task model
</span></span><span class="line"><span class="cl">→ calibration of goals and boundaries
</span></span><span class="line"><span class="cl">→ authorized plan and action
</span></span><span class="line"><span class="cl">→ observation of consequences
</span></span><span class="line"><span class="cl">→ human evaluation and system revision
</span></span><span class="line"><span class="cl">→ next cycle
</span></span></code></pre></div><blockquote>
<p><strong>Effective human–AI collaboration is a continuing process of alignment, action, observation, and correction around a shared task.</strong></p>
</blockquote>
<h2 id="intentions-are-not-fully-formed-before-language">Intentions Are Not Fully Formed Before Language</h2>
<p>A person does not always begin with a complete objective waiting to be encoded into a perfect prompt. Intentions often become clearer through articulation, comparison, and trial.</p>
<p>A request can contain several levels:</p>
<table>
  <thead>
      <tr>
          <th>Level</th>
          <th>Question</th>
          <th>Example</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Utterance</td>
          <td>What was said?</td>
          <td>“Handle this article.”</td>
      </tr>
      <tr>
          <td>Operational intent</td>
          <td>What should the AI do?</td>
          <td>Summarize, edit, rewrite, or publish?</td>
      </tr>
      <tr>
          <td>Task purpose</td>
          <td>Why do it?</td>
          <td>Public release, internal review, or private understanding?</td>
      </tr>
      <tr>
          <td>Value boundary</td>
          <td>What outcomes are acceptable?</td>
          <td>Preserve the argument, protect privacy, verify claims</td>
      </tr>
  </tbody>
</table>
<p>Only part of this structure is usually explicit. The rest may live in earlier turns, the active document, established conventions, institutional rules, or judgments the person has not yet made.</p>
<p>Longer prompts can supply more evidence. They cannot eliminate the underlying problem. Length can also add contradiction, noise, and false precision.</p>
<h2 id="the-system-models-a-task-not-a-whole-person">The System Models a Task, Not a Whole Person</h2>
<p>An AI system has access to evidence such as:</p>
<ul>
<li>the current wording;</li>
<li>conversation history;</li>
<li>visible files and interfaces;</li>
<li>tool observations;</li>
<li>stored preferences and rules;</li>
<li>the user&rsquo;s acceptance or correction of intermediate results.</li>
</ul>
<p>From these signals it constructs a working task model:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">What is the object?
</span></span><span class="line"><span class="cl">What change is requested?
</span></span><span class="line"><span class="cl">What constraints apply?
</span></span><span class="line"><span class="cl">Which facts remain unknown?
</span></span><span class="line"><span class="cl">Which actions are authorized?
</span></span><span class="line"><span class="cl">What observable state counts as completion?
</span></span></code></pre></div><p>That model can become highly accurate without becoming a complete representation of the user&rsquo;s mind. Restricting claims to available evidence prevents the system from presenting speculation about a person as fact.</p>
<h2 id="three-kinds-of-understanding">Three Kinds of Understanding</h2>
<h3 id="semantic-understanding">Semantic understanding</h3>
<p>Can the system resolve the language? For example, what does “use the first one” refer to in the preceding exchange?</p>
<h3 id="operational-understanding">Operational understanding</h3>
<p>Can it turn the language into a concrete task? Which article, language, format, repository, and action are involved?</p>
<h3 id="outcome-understanding">Outcome understanding</h3>
<p>Does it know what state would satisfy the purpose? Is a generated file enough, or must the work be committed, deployed, and verified at a public URL?</p>
<p>These levels can separate:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">understanding the sentence
</span></span><span class="line"><span class="cl">≠
</span></span><span class="line"><span class="cl">knowing what operation to perform
</span></span><span class="line"><span class="cl">≠
</span></span><span class="line"><span class="cl">knowing what completion looks like
</span></span></code></pre></div><p>Many failures arise from disagreement about objects, permissions, or completion evidence even when the words were parsed correctly.</p>
<h2 id="building-a-shared-task-model">Building a Shared Task Model</h2>
<p>Human and system gradually establish a shared, revisable representation of the task. It normally includes:</p>
<table>
  <thead>
      <tr>
          <th>Element</th>
          <th>Question</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Object</td>
          <td>Which file, page, account, product, or problem is being changed?</td>
      </tr>
      <tr>
          <td>Goal</td>
          <td>What change should occur?</td>
      </tr>
      <tr>
          <td>Constraints</td>
          <td>Which facts, formats, styles, privacy limits, and rules must hold?</td>
      </tr>
      <tr>
          <td>Evidence</td>
          <td>What is known, inferred, disputed, or unavailable?</td>
      </tr>
      <tr>
          <td>Authorization</td>
          <td>How far may the system act?</td>
      </tr>
      <tr>
          <td>Completion</td>
          <td>What observable result establishes success?</td>
      </tr>
  </tbody>
</table>
<p>This does not require the person to specify everything at once. A system can use established context and produce reversible work while seeking information only where it changes the result or the permission boundary.</p>
<p>Good collaboration retains decisions already made. It also remains open to revision when new evidence conflicts with an earlier interpretation.</p>
<h2 id="clarify-assume-or-act">Clarify, Assume, or Act?</h2>
<p>Ambiguity does not force a choice between blind guessing and endless questioning. The practical rule depends on:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">degree of ambiguity × cost of error × reversibility
</span></span></code></pre></div><h3 id="low-cost-and-reversible">Low cost and reversible</h3>
<p>Proceed with a stated assumption or a draft. The result itself can help the person clarify the target.</p>
<h3 id="moderate-ambiguity">Moderate ambiguity</h3>
<p>Preserve alternatives, complete the common work, or present comparable options.</p>
<h3 id="high-consequence-or-difficult-to-reverse">High consequence or difficult to reverse</h3>
<p>Confirm the object, scope, and authorization before publication, payment, external messaging, or irreversible deletion.</p>
<p>The purpose of clarification is to control consequences. It should occur where a distinction materially changes the action.</p>
<h2 id="a-prompt-is-evidence-not-a-complete-contract">A Prompt Is Evidence, Not a Complete Contract</h2>
<p>A prompt directly constrains the current task, but its force still depends on context and conversational commitments.</p>
<p>Closely related sentences license different actions:</p>
<ul>
<li>“Can this be published?” requests an assessment.</li>
<li>“Prepare this for publication” authorizes editing.</li>
<li>“Publish this on the site” authorizes an external action.</li>
<li>“I may publish this later” reports a possibility.</li>
</ul>
<p>A system must distinguish questions, proposals, background information, corrections, and authorization.</p>
<p>Conversation also creates durable commitments. Once the person selects an option, approves publication, or defines a format, those decisions should guide later steps. The operative instruction is distributed across the interaction, not confined to the latest sentence.</p>
<h2 id="action-tests-understanding">Action Tests Understanding</h2>
<p>Restating a request does not prove that both sides share the same model. Action exposes hidden disagreement.</p>
<p>A person says, “Update the article on the website.” The system may create local Markdown but fail to commit it; commit without deployment; deploy one language but omit the other; publish both pages but leave discovery files stale.</p>
<p>Verification must therefore proceed through layers:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">task interpretation is correct
</span></span><span class="line"><span class="cl">→ artifact content is correct
</span></span><span class="line"><span class="cl">→ action actually occurred
</span></span><span class="line"><span class="cl">→ external state changed
</span></span><span class="line"><span class="cl">→ the outcome satisfies the purpose
</span></span></code></pre></div><p>Each layer produces evidence. A failed action also reveals which assumption or completion criterion was missing.</p>
<h2 id="the-joint-action-loop">The Joint Action Loop</h2>
<p>The process can be described in eight stages.</p>
<h3 id="1-the-person-encounters-a-problem">1. The person encounters a problem</h3>
<p>A need, obstacle, opportunity, or unsatisfactory result creates pressure to change the current state.</p>
<h3 id="2-a-provisional-intention-forms">2. A provisional intention forms</h3>
<p>The person selects a direction, although the goal, method, and standard may remain incomplete.</p>
<h3 id="3-the-request-is-externalized">3. The request is externalized</h3>
<p>Language carries part of the intention into a prompt, together with available context and constraints.</p>
<h3 id="4-the-ai-constructs-candidate-interpretations">4. The AI constructs candidate interpretations</h3>
<p>The system identifies objects, actions, constraints, missing information, and competing task models.</p>
<h3 id="5-the-task-is-calibrated">5. The task is calibrated</h3>
<p>History, paraphrase, drafts, options, or a necessary question make the interpretation clear enough for the next action.</p>
<h3 id="6-the-ai-acts-within-authorization">6. The AI acts within authorization</h3>
<p>The system plans steps, invokes tools, preserves state, and checks permission at consequential boundaries.</p>
<h3 id="7-consequences-are-observed">7. Consequences are observed</h3>
<p>Files, command results, public pages, user reactions, and other evidence show what actually happened.</p>
<h3 id="8-both-sides-revise">8. Both sides revise</h3>
<p>The person may change the goal, correct a mismatch, or accept the result. The system updates its task model and continues or stops.</p>
<p>The cycle is:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">intention
</span></span><span class="line"><span class="cl">→ expression
</span></span><span class="line"><span class="cl">→ interpretation
</span></span><span class="line"><span class="cl">→ calibration
</span></span><span class="line"><span class="cl">→ action
</span></span><span class="line"><span class="cl">→ observation
</span></span><span class="line"><span class="cl">→ evaluation
</span></span><span class="line"><span class="cl">→ revised intention
</span></span></code></pre></div><h2 id="a-division-of-responsibility">A Division of Responsibility</h2>
<p>A joint loop does not make human and system responsibility identical.</p>
<h3 id="the-person-contributes">The person contributes</h3>
<ul>
<li>the real situation that needs to change;</li>
<li>value judgments and ultimate purpose;</li>
<li>private context and unspoken constraints;</li>
<li>authorization for consequential actions;</li>
<li>final acceptance of whether the outcome is worth having.</li>
</ul>
<h3 id="the-ai-system-contributes">The AI system contributes</h3>
<ul>
<li>a structured interpretation of available evidence;</li>
<li>alternative plans and their relevant differences;</li>
<li>executable steps and tool use;</li>
<li>explicit uncertainty, permission, and completion states;</li>
<li>rapid revision when new evidence arrives.</li>
</ul>
<h3 id="the-platform-or-organization-contributes">The platform or organization contributes</h3>
<ul>
<li>identity and access control;</li>
<li>data and privacy boundaries;</li>
<li>logs, versioning, and recovery;</li>
<li>failure handling and assignment of accountability;</li>
<li>observable status for external actions.</li>
</ul>
<p>The AI cannot settle every value question for a person. A person cannot explain every failure by pointing only to an isolated model. Outcomes belong to the complete sociotechnical arrangement.</p>
<h2 id="where-the-loop-breaks">Where the Loop Breaks</h2>
<h3 id="a-feeling-is-mistaken-for-a-complete-objective">A feeling is mistaken for a complete objective</h3>
<p>A person knows that an output is wrong but cannot yet specify the desired alternative. The system should help compare concrete possibilities rather than assume a unique hidden answer.</p>
<h3 id="the-most-likely-interpretation-becomes-true-intent">The most likely interpretation becomes “true intent”</h3>
<p>High probability means that available evidence favors an interpretation. It does not reveal the person&rsquo;s complete private purpose.</p>
<h3 id="content-approval-becomes-action-approval">Content approval becomes action approval</h3>
<p>Accepting an article does not automatically authorize its public release.</p>
<h3 id="a-plan-is-reported-as-completion">A plan is reported as completion</h3>
<p>“We will update the site” is not deployment evidence. Local files, commits, deployments, and live pages are different states.</p>
<h3 id="only-the-output-is-evaluated">Only the output is evaluated</h3>
<p>A correct answer can result from an unreliable method. A failed attempt can expose an important unknown. Durable collaboration evaluates results, evidence, and the capacity to correct.</p>
<h3 id="feedback-does-not-update-the-next-cycle">Feedback does not update the next cycle</h3>
<p>If corrections are neither retained nor applied, the same mismatch repeats and no learning loop forms.</p>
<h2 id="improving-the-loop">Improving the Loop</h2>
<p>A person can improve collaboration by:</p>
<ul>
<li>describing the desired change, not only naming an operation;</li>
<li>distinguishing exploration, drafting, revision, approval, and publication;</li>
<li>stating unacceptable outcomes at consequential points;</li>
<li>locating feedback at the level that was misunderstood;</li>
<li>allowing the intention itself to change after new evidence.</li>
</ul>
<p>An AI system can improve collaboration by:</p>
<ul>
<li>separating explicit requirements, inferences, and unknowns;</li>
<li>carrying forward confirmed context;</li>
<li>using reversible artifacts to advance low-risk work;</li>
<li>checking object and authorization at high-consequence boundaries;</li>
<li>reporting observed results instead of plans;</li>
<li>updating its task model after correction.</li>
</ul>
<h2 id="conclusion">Conclusion</h2>
<p>A human encounters a situation, develops needs and intentions, and expresses only part of them in language. An AI system uses visible evidence to construct a task model, reason about options, and act through tools. The consequences and the person&rsquo;s evaluation then revise that model and may revise the original intention.</p>
<p>The central object is therefore not a perfect one-shot prompt. It is a shared loop that remains correctable:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">human purpose, values, and authorization
</span></span><span class="line"><span class="cl">+
</span></span><span class="line"><span class="cl">AI interpretation, planning, and execution
</span></span><span class="line"><span class="cl">+
</span></span><span class="line"><span class="cl">environmental consequences and evidence
</span></span><span class="line"><span class="cl">+
</span></span><span class="line"><span class="cl">feedback that changes the next cycle
</span></span></code></pre></div><blockquote>
<p><strong>Reliable collaboration does not require an AI to read an invisible “true mind.” It requires both sides to make the current task, action boundary, and evidence of completion progressively clearer.</strong></p>
</blockquote>
<h2 id="further-reading">Further Reading</h2>
<ul>
<li><a href="/en/notes/ai-user-intent-inference/">User Intent in AI: Inference Under Uncertainty</a></li>
<li><a href="/en/notes/ai-reasoning-and-action/">How AI Systems Reason and Act</a></li>
<li><a href="/en/notes/human-thinking/">How Human Thinking Works: Representation, Reasoning, and Action</a></li>
<li><a href="https://aclanthology.org/2024.emnlp-main.119/">EMNLP 2024: Making Language Models Explicitly Handle Ambiguity</a></li>
</ul>
]]></content:encoded></item></channel></rss>