Entities: Identity, Reference, and AI Systems

An entity is something treated as having a distinguishable identity. This essay separates mentions, names, classes, records, and real-world referents across philosophy and AI.

People, companies, cities, products, documents, events, and fictional characters are very different kinds of things. Yet an information system may treat all of them as entities.

The word is broad because it does not describe a particular material composition. It describes a role within thought, language, or a model:

An entity is something treated as having a distinguishable identity, so that it can be referred to, described, related to other things, tracked over time, or acted upon.

An entity need not be tangible. It need not exist independently. It need not even exist in the actual world. What it needs, within a particular domain, is enough identity for the system to treat it as the same subject across multiple statements or operations.

Where does the word “entity” come from?

English entity comes through Medieval Latin entitas, formed from ens, a being or something that is, which in turn is related to the Latin verb esse, to be. The word therefore carries an ontological background: it concerns something considered as a being or item of existence.American Heritage Dictionary: entity

That history explains why entity can be used so widely. It may refer to a physical object, a person, an institution, an event, a number, a proposition, or another item admitted by a theory. In philosophy, thing, being, entity, and object may all compete for the role of a maximally general term for whatever a system acknowledges.Stanford Encyclopedia of Philosophy: Object

No single list of entities is philosophically neutral. A physicalist ontology, a mathematical ontology, a legal ontology, and a fictional world may recognize different kinds of things. Calling something an entity is therefore both a semantic move and, in many contexts, an ontological commitment.

An entity is not merely a physical object

A physical object usually has material structure and some spatial boundary. An entity need not.

A corporation has no single body. It depends on law, records, roles, property, contracts, and continued institutional recognition. A meeting is not a durable object, but it can have participants, a time, a location, an agenda, and an outcome. An account exists only within a platform and its rules, yet the platform must still distinguish one account from another.

All three can function as entities because each can be identified and become the subject of further claims.

The English words entity and substance should also be separated. An entity is any item treated as a being or object of reference. A substance, in major philosophical traditions, is more specifically something taken to exist relatively independently or to bear properties. Events, relations, numbers, and properties may count as entities without counting as substances in that stronger sense.Stanford Encyclopedia of Philosophy: Substance

Reference does not establish real existence

Sherlock Holmes can be named, described, compared with other characters, and placed in a network of fictional relations. He is an entity in literary discourse and may be an entity in a knowledge base. None of this makes him a historical person.

At least three questions must therefore remain separate:

LevelQuestion
Discourse entityHas language introduced something that can be referred to again?
Model entityDoes an information system represent it as a distinct item?
Real-world entityIs there sufficient evidence for a corresponding thing in the actual world?

An entity can exist at the first two levels without satisfying the third. Plans, hypothetical products, possible events, mistaken identities, and fictional characters all demonstrate why representation and reality must not be collapsed.

Identity is the central problem

Finding a noun is easy. Determining what makes something the same entity is harder.

A person may change names and addresses while remaining the same person. A corporation may replace every employee and continue as the same legal organization. An article may be revised many times while retaining one publication history. A product may keep its commercial name while its capabilities, components, or terms change substantially.

Different types of entities require different identity criteria:

Entity typePossible basis of identity
Personlegal records, bodily and biographical continuity
Corporationlegal registration, organizational and contractual continuity
Documentidentifier, provenance, content hash, or version history
Articleauthorship and publication lineage, slug, revision history
Product modelmodel definition, capability boundary, specification
Commercial itemSKU, serial number, batch, or unit of sale
Online accountplatform, stable account ID, control and authentication
Eventparticipants, time, place, and occurrence structure

A name is not an identity. Two people can share a name, and one person can use several names. A set of properties is not automatically an identity either. Properties change, records conflict, and two objects may resemble one another closely without being the same object.

A useful entity representation usually needs more than a label:

stable identifier
+ entity type
+ attributes
+ relationships
+ temporal state
+ provenance
+ confidence in the identity match

How does ontology relate to entities?

An entity does not become a useful modeling unit in isolation. A system must decide what kinds of entities it recognizes, which attributes they may have, which relations may connect them, and what counts as persistence or change.

Those decisions form part of an ontology.

An entity is a particular item recognized within a domain; an ontology states which kinds of items the domain recognizes and how they may be organized.

“Shanghai” may be stored as a string in a simple customer table. In a geographical knowledge system, it may be represented as an entity with coordinates, administrative status, districts, and relationships to other places. The difference depends on whether the system needs to identify Shanghai independently and reason about it.

Philosophical ontology asks what exists and how different kinds of beings exist. Computational ontology turns a domain commitment into an explicit model of classes, entities, properties, relations, and constraints. A fuller account appears in “What Is Ontology? From What Exists to Models AI Can Use”. Here ontology matters because it supplies the type system and identity conditions within which entities can be distinguished.

What does “entity” mean in AI?

In AI, an entity is usually an object distinguished from text, images, records, or an environment because the system needs to understand, retrieve, remember, reason about, or act on it.

The term changes meaning across tasks. Named-entity recognition, entity linking, entity resolution, knowledge graphs, computer vision, databases, and AI agents do not operate at exactly the same level.

Named-entity recognition finds mentions, not verified objects

Consider the sentence:

Apple plans to announce a new phone in Shanghai tomorrow.

A named-entity recognition system may label:

  • Apple as an organization;
  • Shanghai as a location;
  • tomorrow as a date.

At this stage, the system has identified mentions: spans of language that appear to refer to named or otherwise categorized entities. In engineering terms, a conventional NER component predicts labeled token spans. spaCy, for example, describes its entity recognizer as identifying non-overlapping labeled spans.spaCy: EntityRecognizer

The output does not yet prove that Apple refers to Apple Inc. rather than a different organization, a title, or an annotation mistake. Nor does it establish that the announced phone is a particular product with a known identity.

The distinction is fundamental:

entity mention ≠ entity identity
entity name ≠ unique referent
entity type ≠ proof of existence

Entity linking connects a mention to a canonical entity

Entity linking attempts to determine which known entity a mention refers to.

“Apple” in a sentence
        ↓ disambiguation
Apple Inc.
        ↓ normalization
knowledge-base ID: company/apple-inc

This requires at least two forms of reasoning:

  • Disambiguation: Which entity with this name fits the context?
  • Coreference and normalization: Do Apple, Apple Inc., and the Cupertino company refer to the same entity here?

Linking converts a piece of language into an addressable object in a knowledge system. It is still fallible. The selected knowledge-base entry may be wrong, duplicated, incomplete, or out of date.

Entity resolution asks whether records describe the same thing

Entity linking usually begins with language and a knowledge base. Entity resolution often begins with records.

A customer system may contain:

Li Ming, Shanghai, 138****1234
Ming Li, 上海市, 138****1234
李明, Pudong, 138****1234

Do these records describe one person, two people, or three? Shared fields provide evidence, but they do not make the answer automatic. Phone numbers can be reassigned, addresses can be shared, and names can collide.

Entity resolution may involve:

  • deduplicating records;
  • matching aliases and transliterations;
  • detecting that one entity has split or merged in a source system;
  • preserving conflicting claims rather than forcing a premature merge;
  • recording why two records were considered the same.

This is an identity decision under uncertainty. A useful system keeps the evidence and confidence behind the match instead of treating similarity as certainty.

Knowledge graphs place entities in a network of claims

In a knowledge graph, entities are commonly represented as nodes connected by typed relations:

Apple Inc. ──headquartered in──> Cupertino
Apple Inc. ──released──> iPhone
iPhone ──instance of──> smartphone product

This representation distinguishes:

  • entities such as Apple Inc., Cupertino, and an iPhone model;
  • classes such as company, city, and product;
  • attributes such as dates and names;
  • relations such as headquartered in and released.

RDF represents claims as subject–predicate–object triples. Its notion of a resource is deliberately broad: a resource may denote a physical thing, a document, an abstract concept, or another item in the universe of discourse.W3C: RDF 1.1 Concepts and Abstract Syntax

An ontology supplies general rules, while a knowledge graph contains particular claims:

ontology: a company may release a product
claim: Apple Inc. released a particular iPhone model

A graph node is not automatically a verified real-world object. It may represent a class, a fictional entity, a planned object, an uncertain hypothesis, or an erroneous record. Provenance and status remain necessary.

Concepts, classes, entities, identifiers, and records

Several layers are easily confused:

LayerWhat does it provide?
Word or phrasean expression used in language
Namea conventional way to refer to something
Concepta general structure used to understand things
Classa modeled category of possible members
Entitythe particular item currently referred to
Identifiera stable handle used by a system
Recordstored claims about an entity
Real-world referentwhatever, if anything, exists beyond the model

Company may be a concept and a class. Apple Inc. may be an entity. Apple may be a name or mention. company_001 may be an internal identifier. A row in a database may contain claims about the company. None of these layers is identical to the organization itself.

One entity may have many names and records. One record may accidentally combine facts about several entities. A unique database key guarantees uniqueness inside a table; it does not prove that the row corresponds correctly to one real-world thing.

When should something be modeled as an entity?

Not every noun phrase needs its own entity. Six questions help:

  1. Does it need a distinct identity? Must the system distinguish this item from similar items?
  2. Will it be referred to repeatedly? Will multiple documents, records, or tasks mention it?
  3. Does it have its own attributes? Must the system store a status, date, location, or owner?
  4. Does it participate in relationships? Must it be connected to other objects?
  5. Does it persist through time? Must the system recognize it after some properties change?
  6. Can the system act on it? Will it be queried, edited, sent, authorized, purchased, deleted, or monitored?

If most answers are yes, an entity is often appropriate. Otherwise a literal value, label, or temporary span may be enough.

Entity modeling is not free. Every entity type introduces identity rules, lifecycle questions, merge and split behavior, provenance requirements, and access-control consequences.

Can events, states, and intentions become entities?

Grammatical categories do not map directly onto model categories. Nouns do not always denote entities, and verbs do not always remain mere relations.

“Maya signed Contract C36” can be represented as a simple relation:

Maya ──signed──> Contract C36

If the system must record the date, location, version, witnesses, method, and legal status, the signing can become an event entity:

Signing Event E1024
├── signer: Maya
├── object: Contract C36
├── time: 2026-09-24
├── place: Shanghai
└── status: effective

An intention can likewise be modeled as a mental-state entity when the system needs to record whose intention it is, what outcome it concerns, when it was inferred, which evidence supports the inference, and how uncertain the interpretation remains.

Reification makes an event, relation, or state available for further description. It is useful when the system needs that detail, but it should not be mistaken for a discovery that the world naturally divides itself in exactly that form.

How do large language models handle entities?

A standard large language model receives tokens and computes distributed representations shaped by training data and current context. It can learn that Apple often occurs near company, product, iPhone, and Cupertino, and it can frequently disambiguate the word from the fruit.

This competence does not imply that the model contains a single, explicit, canonical Apple Inc. record comparable to a carefully maintained knowledge base. Entity information may be distributed across parameters and reconstructed probabilistically in context.

Consequently, a language model may:

  • resolve an entity correctly in one context and confuse it in another;
  • conflate a company, its brand, its products, and its website;
  • recall an old property without knowing that it has changed;
  • invent a plausible person, paper, organization, or product;
  • answer without a stable source for the entity claim.

For tasks that require reliable action, model output is usually only one layer:

language model
+ domain ontology
+ entity IDs and resolution
+ current state and provenance
+ time-aware records
+ permissions and action rules

The language model interprets open-ended language. The ontology defines possible types and relations. Entity services determine identity. Data sources establish current state. Authorization determines which operations are permitted.

Why do AI agents need explicit entity identity?

Consider the instruction:

Update yesterday’s article about needs on Moonment.

The request contains several unresolved references:

ExpressionEntity question
yesterdayWhich timezone and time interval?
the article about needsWhich document and slug?
updateEdit a draft, commit a repository, or publish a website?
MoonmentWhich project, repository, deployment, and domain?
article versionChinese, English, or both?
requesting userWhich permissions and prior authorization apply?

The agent may understand the general intention while still acting on the wrong file, project, account, language version, or deployment.

A safer path is:

natural-language request
        ↓
mention and reference detection
        ↓
type assignment and coreference
        ↓
entity linking and resolution
        ↓
relation, event, and intent interpretation
        ↓
state, provenance, and authorization checks
        ↓
action on the resolved entity

Intent specifies the change the user appears to seek. Entity resolution determines what the intended action applies to. Authorization determines whether the action may be performed. None can substitute for the others.

Can AI establish that an entity really exists?

Language alone cannot establish existence.

An AI system can estimate that a phrase is probably a person’s name, that context probably indicates a company, or that a mention probably links to a known record. Real-world verification requires additional evidence, such as:

  • an authoritative registry or primary source;
  • a current database record;
  • a file that actually exists in the relevant filesystem;
  • a live website or API;
  • an authenticated identity and permission system;
  • a sensor observation or human confirmation.

The following claims are distinct:

the text mentions something
≠ the model identified it correctly
≠ the system linked the correct record
≠ the record is accurate and current
≠ the corresponding thing exists now
≠ the user is authorized to act on it

Model confidence is not proof of existence. It describes a model’s judgment under particular inputs and assumptions. It does not replace provenance or verification.

A checklist for entity design in AI systems

  1. Which entity types does the system recognize, and why?
  2. Are mentions, names, classes, entities, identifiers, and records kept separate?
  3. What establishes identity for each entity type?
  4. How are aliases, namesakes, renaming, and duplicate records handled?
  5. Do attributes and relations carry time and provenance?
  6. Are events and changing states being flattened into misleading static fields?
  7. Are extraction, linking, resolution, and real-world verification separate stages?
  8. Can the system distinguish fictional, planned, hypothetical, and actual entities?
  9. What happens to historical claims when entities merge, split, or change type?
  10. Before acting, how does the system verify identity, current state, and permission?

An entity-rich system can still be unreliable. Without identity criteria, provenance, temporal state, and authorization, it merely attaches confident-looking labels to uncertain referents.

Entities connect language to action

Entities form an interface between language, knowledge, and operations:

objects, events, and states in a domain
                ↓
names and descriptions in language
                ↓
entity, attribute, and relation recognition
                ↓
links to records and external evidence
                ↓
interpretation of needs and intentions
                ↓
authorized action and recorded outcomes

An entity answers, “Which particular thing are we talking about?” An attribute answers, “What is it like?” A relation answers, “How is it connected?” An event answers, “What happened?” An intention answers, “What outcome is an agent trying to bring about?”

The most serious entity error in AI is not failing to label a noun. It is treating an ambiguous expression as though it already named one verified, current, and actionable object.

References

If this was useful, subscribe via RSS.

Content is open for citation with attribution; please link back to the source.