Identity and Access Substrate
Every identity and access control resolves against the link between a subject and the identity representing it, and almost no enterprise measures how strong that link is. This report defines the substrate vocabulary and grades binding confidence across every identity constituency.
Observations
These observations describe the state of the identity and access substrate as CAI finds it. Each is empirical and each has a control consequence.
- Enterprises can name their identity capabilities and cannot state how well their identities are bound. A program can produce an inventory of its identity provider, its governance platform, and its privileged access tooling. Ask the same program what proportion of its identity population is linked to a real subject on the strength of evidence rather than the strength of an assertion, and the question is usually reclassified as a data quality matter and routed to whoever runs the directory. It is not a data quality matter. Every control the program operates resolves against that link.
- The identity assurance framework the industry defers to covers natural persons only. 1 The identity assurance levels that enterprises cite in policy, that vendors certify against, and that regulators reference are scoped by their authors to human beings. Machines, workloads, agents, and organizations sit outside that scope by design rather than by oversight. The majority constituency in a modern estate therefore has no assurance ladder at all, and the absence goes unremarked because the conversation about machine identity is a conversation about counting, not about confidence.
- The industry cannot agree how many machine identities exist, and the disagreement is a substrate symptom rather than a measurement problem. 2 Reputable sources differ by an order of magnitude. Some count credentials, some count accounts, some count service principals, some count running workloads. A population cannot be counted before it has been defined, and the counts are published anyway. In CAI’s assessment the divergence is not a failure of measurement. It is direct evidence that the substrate vocabulary is unsettled.
- Ownership of machine subjects survives in tickets rather than in records. Service accounts and workloads are created at the moment they are needed, by whoever needs them. The accountable human is captured in a change request or an approval thread, each of which decays faster than the account does. When the creator leaves, the binding does not fail loudly. It quietly stops being true, and in most cases nothing in the estate registers the change.
- Resource and environment vocabulary differs inside a single enterprise. Two teams in the same organization will describe one protected object as an application, a service, a data set, and an integration. Coverage claims built on that vocabulary are not comparable between teams, between an enterprise and its auditor, or between an enterprise and a vendor’s account of what it covers.
Positions
CAI takes five positions on the substrate. Each is stated as an action an enterprise or a vendor can take.
- Treat identity binding as a control surface with a named owner. The link between a subject and the digital identity representing it is not a byproduct of directory hygiene. It is the assumption on which every governance decision, every runtime decision, and every attribution rests. Give it an owner, a definition of adequacy, and a measurement, in the same way an enterprise does for any other control.
- Grade binding confidence and refuse the binary. An identity is not verified or unverified. It is bound at a level, and the level is a property of a class of identity rather than of an enterprise as a whole. A program that reports a single verification percentage has discarded the information that would let it act.
- Establish the account as an object in its own right. The account is not a view of the identity and the identity is not inferable from the account. Each has a lifecycle, an owner, and a set of evidence. Estates that conflate the two discover orphaned and dormant accounts by scanning, because they have no record that would have told them.
- Adopt one subject taxonomy across human, machine, workload, and agent constituencies before the agent population sets the convention by default. Agents are being onboarded now, into estates whose taxonomy already fails for service accounts. The window in which a single taxonomy remains reachable is measured in a small number of years, and it closes by accretion, not by decision.
- Treat the operating environment as a determinant of control rather than as deployment detail. The environment an access sits in decides what is protectable, what context is available at the moment of decision, and what granularity of enforcement is achievable. It belongs on the coverage axis for that reason and not because it is convenient bookkeeping.
The Control and Coverage Matrix in Brief
This report is one of a set that share a single organizing instrument, the CAI Control and Coverage Matrix. The matrix is described in full in the Root Document; the summary here is the canonical short form carried by every Foundation Report so that each reads standalone.
The matrix has a control axis, a coverage axis, and a substrate beneath both. The control axis has three durable categories.
- Governance is what happens before access: administration, entitlement, certification, and policy.
- Runtime is what happens at the moment of access: authentication, authorization, session, federation, and enforcement.
- Observability is what happens after and around access: detection, posture, signals, and analytics that feed back into the other two.
This report treats the substrate beneath both axes.
The coverage axis has three facets that qualify every access at once.
- Identity constituency: human (workforce, customer, partner), machine (device, workload, agent), or organization.
- Technical environment: channel, infrastructure, data, and application, each across on-premises and cloud.
- Business environment: operating unit, subsidiary, and jurisdiction.
Beneath both axes sits the substrate every control depends on: the identity and access data on which decisions are made, and the identity verification and assurance by which a subject is established.
Each intersection of a control category with a coverage area is a control-and-coverage instance: a discrete unit of accountable work, delivered by a capability, operated by someone, and owned by a named person. An enterprise’s surface is the full set of instances that apply to it; its posture is how completely those instances are covered and owned. Capabilities are the unit of delivery, and one capability typically serves several instances. Instances are the unit of accountability, and each has exactly one accountable owner.
CAI View The substrate is the only element of the matrix that no control category owns and every control category consumes. That is why it is the last thing an enterprise inspects and the first thing that explains why its controls underperform.
What the Substrate Is, and What Rests On It
The substrate is the set of facts about subjects, identities, accounts, resources, and environments that every identity and access control assumes to be settled before it operates. Governance assumes it knows who is being governed. Runtime assumes it knows who is asking. Observability assumes it knows to whom an event is attributable. None of the three establishes those facts. Each inherits them.
That inheritance is the whole point. No control can be more correct than the facts it resolves against. An entitlement certification performed against an identity that represents two people is not a weak certification, it is a certification of the wrong thing. An authorization decision made for a workload whose identifier is shared by forty instances is not a granular decision, it is a decision at the granularity of the identifier rather than the granularity anyone intended. An attribution to an account whose owner departed eighteen months ago is not a low-confidence attribution, it is an attribution to nobody.
This is why the substrate is not itself a piece of infrastructure. Identity and access is infrastructure, and CAI treats it as such elsewhere in this series. But every element of that infrastructure carries two independent decisions: how it is realized, by building, implementing, or blending, and who operates it, the enterprise or a vendor. The substrate carries neither. Nobody builds it, nobody implements it, and nobody runs it. It is the condition left behind by how everything else was realized and operated, and an enterprise is in that condition whether or not it has ever inspected it.
Table 1. The substrate objects, and which control category consumes each
| Substrate object | What it establishes | Which controls inherit it |
|---|---|---|
| Subject | That there is a real party, of a known kind, on whose behalf access occurs | All three. Governance governs it, runtime authenticates it, observability attributes to it |
| Digital identity | That the subject has a representation the enterprise can act on | All three |
| Identifier | That the representation can be named and correlated across systems | Observability first, then governance |
| Account | That the identity can act inside a specific target system | Governance holds entitlements here; observability sees actions here |
| Resource | That there is a defined object an action operates on | Runtime enforces at it; governance grants toward it |
| Operating environment | That the context of the access is known, and with it what is enforceable | Runtime and observability directly, governance indirectly |
| Identity and access data | That the above are recorded somewhere trusted enough to decide against | All three, continuously |
| Verification and assurance | That the link between subject and identity rests on something | All three, usually without inspecting it |
Why the substrate stays invisible
Three mechanisms keep the substrate out of view, and they reinforce one another.
- Controls report on themselves, not on what they resolved against. A certification campaign reports completion rates. An authorization engine reports decision latency and policy coverage. A detection pipeline reports alert volume and mean time to triage. None of these instruments has a field for the quality of the identity the work was performed on, so a program with an excellent dashboard can be operating almost entirely on assertions.
- The substrate degrades without an event. Most control failures produce a signal. Substrate failures produce silence. An employee leaves and a service account he or she created continues working perfectly. The binding became false at the moment of departure, and nothing observed the transition, because nothing was watching a fact that no system owned.
- Responsibility is distributed to the point of disappearance. Human resources owns employment facts. Engineering owns workload creation. Procurement owns partner relationships. The identity team owns the directory. Each owns a fragment and none owns the binding, so when the binding is wrong there is no one whose job it was.
The combined effect is that substrate weakness is discovered during incidents, audits, and migrations, which are the three worst moments to discover it and the three occasions on which remediation is most expensive.
How the substrate becomes visible
The substrate becomes inspectable when three questions can be answered from records rather than from recollection.
- Who is the subject behind this identity.
- What did that answer rest on when it was established.
- When was it last confirmed.
An enterprise that can answer all three for a class of identity can state that class’s binding confidence. An enterprise that can answer none of them is not at a low band. It does not know which band it is at, which is a different and worse position.
These questions are answerable from data most enterprises already hold, which is the practical reason substrate assessment is tractable. The information is generally present and scattered, not absent. What is missing is the instrument that says what to do with it.
Watch-Out A weak substrate does not present as a substrate problem. It presents as unreliable certifications, noisy detections, and authorization that cannot be made granular. Programs then buy against the symptom.
Boundaries and the Handoffs
This foundation report defines the substrate. It does not operate on it. The distinction determines what belongs in this report and what belongs elsewhere in the Body of Knowledge.
Table 2. What this report supplies, and what it hands off
| This report supplies | This report hands off |
|---|---|
| The definition of subject, identity, identifier, account, resource, and operating environment, and the relationships between them | The lifecycle that creates and retires them, to the Identity and Access Governance Controls Foundation Report |
| The definition of identity verification and of assurance as properties of the substrate | Proofing methods, evidence types, and the vendor landscape, to the Identity Proofing and Verification Capability Report |
| Binding confidence as a graded property, and the bands that express it | The credential model and what happens when a credential is presented, to the Identity and Access Runtime Controls Foundation Report |
| The definition of identity and access data and of identifier space | Telemetry, attribution, and the operational data foundation, to the Identity and Access Observability Controls Foundation Report |
| The vocabulary in which the coverage axis is expressed | Directory and registry product treatment, to the relevant Capability Report |
The handoffs also run in the other direction, which is easy to miss because the substrate is drawn beneath the control axis rather than beside it.
- Governance returns lifecycle events that should update the substrate, including the departures and transfers that silently invalidate bindings.
- Runtime returns evidence of successful and failed verification that is itself information about binding strength.
- Observability returns the correlation failures that reveal where the identifier space is broken.
Most enterprises have built the outbound handoffs and none of the inbound ones, which is why substrate quality does not improve over time even in programs that are otherwise maturing.
One handoff deserves emphasis because it is regularly collapsed. Verification is an activity. Assurance is a property that results from it. This report treats the property, because the property is what every downstream control inherits. The activity, including how evidence is collected, validated, and checked against the applicant, belongs to the capability layer.
Definitions
Each definition below is the canonical CAI statement of the term. Where a definition differs from a working definition carried in an earlier Foundation Report, the statement here governs. Usage notes appear only where the term is routinely misapplied.
Subject
The real party on whose behalf access occurs. A subject may be a person, an organization, or a machine, where machine covers devices, workloads, and agents. The subject exists independently of any system that records it, and it is not the same thing as the digital identity that represents it, the account through which it acts, or the credential by which it proves possession of that identity. Usage note: the four are routinely spoken of as one, and almost every substrate defect begins there.
Digital identity
The representation of a subject inside a system of record, carrying the attributes the enterprise holds about that subject. The intended cardinality is one identity for one subject. The observed cardinality is many, and the distance between the two is a measurable defect rather than an accepted condition.
Identifier
A value that distinguishes one identity or one account from others within a defined scope. An identifier has meaning only inside the scope that issued it. Two identifiers with the same value, issued by different authorities, may or may not denote the same thing, and nothing in the value itself resolves the question.
Unique identifier
An identifier guaranteed to distinguish one identity from every other within a stated scope, and not reassigned within that scope. Uniqueness is a property of the issuing authority’s discipline, not of the value’s format. A value that looks globally unique and is reassigned on reuse is not a unique identifier.
Identifier space
The set of identifiers by which subjects are named across an estate, together with the rules by which those identifiers resolve to one another. The identifier space is the substrate object that correlation depends on, and it fails through collision, through reuse after deprovisioning, and through the absence of any stable identifier for parts of the machine population.
Account
The object through which an identity acts within one target system, holding that system’s entitlements and carrying its own lifecycle and its own owner. An account is not a projection of an identity and an identity is not inferable from an account. Usage note: accounts are not necessarily persistent. Just-in-time and ephemeral accounts are accounts, and the substrate must describe them without treating persistence as definitional.
Credential
The means by which a subject demonstrates control of a digital identity. The substrate defines what the credential is bound to, and nothing more. The substrate establishes what a credential is bound to. What happens when a credential is presented is a runtime concern and is treated in the Identity and Access Runtime Controls Foundation Report.
Resource
The defined object that a permitted action operates on: an application function, a data object, a channel, or an element of infrastructure. A resource is protectable only to the granularity at which it is defined, which is why resource definition and enforcement granularity are the same problem seen from two directions.
Operating environment
The technical context an access sits in, expressed across channel, infrastructure, data, and application, each spanning on-premises and cloud. Usage note: the environment is one facet of the coverage axis and not the whole of it. The business context in which an access occurs, meaning operating unit, subsidiary, and jurisdiction, is a separate facet and is not part of the operating environment.
Identity and access data
The recorded facts about subjects, identities, identifiers, accounts, entitlements, resources, and environments on which every access decision is made. It is distinguished from the operational telemetry that observability collects about what happened. The two are related, but the first is the basis of decisions and the second is the record of them.
Authoritative source
The system of record trusted to assert a fact about a subject or a resource. Authority is per-fact and per-population, not global. A system may be authoritative for a person’s employment status and carry no authority at all over that same person’s legal name. Usage note: for the workforce there is usually an authoritative source. For several other constituencies there is none, and the substrate should say so rather than nominate a convenient substitute.
Identity verification
The activity of establishing the link between a digital identity and the subject it claims to represent, by validating evidence about that subject. Verification happens at a point in time and produces a property that persists after it, decays after it, or was never established at all.
Identity assurance
The confidence that a digital identity corresponds to the subject it claims to represent. Assurance is the property; verification is the activity that produces it. Assurance is graded, and the grade belongs to a class of identity, not to an enterprise.
Identity binding
The established link between a subject and the digital identity that represents it. The binding is what every downstream control resolves against, and it is the substrate’s unit of control.
Binding confidence
The graded strength of an identity binding, expressed on the S0 to S4 scale defined in this report. Binding confidence is a property of a class of identity within a constituency, and it is stated per class, not as a single figure for an estate.
Applicant and subscriber
An applicant is a subject undergoing the process that would establish a binding. A subscriber is the same subject after the binding has been established and recorded. The distinction matters to the substrate because the evidence that supports the binding is available while the subject is an applicant and is frequently discarded once the subject becomes a subscriber.
Registry and directory
A registry is a system that records the existence and attributes of identities. A directory is a registry optimized for lookup and for serving authentication and authorization. Neither is an authoritative source by virtue of being a registry or a directory; authority is conferred by the facts a system is trusted to assert, not by the software category it belongs to.
Trust domain
The administrative scope within which an identifier has meaning and within which an issuing authority guarantees uniqueness. Identifiers do not carry their trust domain in their value, which is why the same value issued in two domains denotes two different things and why cross-domain correlation requires the relationship to be recorded rather than assumed.
Delegation
The arrangement by which one subject acts on behalf of another. Delegation is a substrate fact rather than a runtime one, because the chain of subjects behind an action must be recorded before any control can evaluate it. Usage note: an agent identity without a recorded delegation chain describes what is acting and not on whose behalf, which is the half of the question that matters.
Identity proofing
The process by which evidence about a subject is collected, validated, and checked against the subject presenting it, producing a binding at a stated level. This report treats the resulting property. The process itself, including evidence types, methods, and the vendor landscape, is treated in the Identity Proofing and Verification Capability Report.
Entitlement
A specific permission held at an account within a target system. Entitlements are named here because they are what governance manages and because they are held at the account rather than at the identity, which is the practical reason the account has to be modeled as an object. Their lifecycle is treated in the Identity and Access Governance Controls Foundation Report.
Identity constituency
The class of subject an identity represents: human, machine, or organization, each with its own subordinate categories. Constituency is the first facet of the coverage axis, and it is the dimension along which binding confidence differs most sharply.
Standards Note The international framework for identity management devotes explicit attention to keeping the terms identity and identifier distinct, and treats that separation as one of its core concepts. 3Enterprises that treat the two as interchangeable are not taking a shortcut. They are discarding the distinction on which correlation depends.
The Conceptual Model
The substrate is a small number of objects and a smaller number of relationships. What matters is the cardinality of each relationship, because that is where intent and reality diverge.

Figure 1. The substrate conceptual model, with the cardinality of each relationship as intended and as observed
Four cardinalities carry the report.
- Subject to digital identity is intended to be one to one and is observed to be one to many. Every enterprise states the intent. No enterprise meets it. What separates programs is whether they measure the gap.
- Digital identity to account is one to many by design. This is not a defect. A person holds accounts in dozens of target systems and should. The defect appears when the estate cannot enumerate which accounts belong to which identity.
- Identity to identifier is one to many across scopes. The same identity is named differently in each system that holds it, and correlation is the work of relating those names. Where no stable identifier exists, correlation is guesswork.
- Workload instance to identifier is many to one, deliberately. 4 The emerging standards for workload identity permit many running instances to share one identifier. Machine binding is therefore coarser than human binding by design, and an enterprise that expects instance-level attribution from a workload identifier has misread the model rather than found a gap.
That last point is worth holding. Much of the frustration enterprises express about machine identity is frustration at a granularity the standards deliberately chose.
Two further relationships carry less weight but cause disproportionate confusion. An identity holds credentials, and the number of credentials tells an observer nothing about the strength of the binding beneath them. An identity bound on nothing can hold a hardware security key and a passkey and remain bound on nothing. And a subject may act on behalf of another subject, which makes the acting party a chain rather than a single node. Estates that model the chain as a single node lose the only information that would let a control evaluate delegated action.
The cardinalities also explain why substrate repair is not a data cleanup project. Cleanup assumes a correct state exists somewhere to be restored. Here the correct state was never captured, because the evidence that would define it was discarded at the point of binding. Repair therefore means establishing bindings again for the classes where the cost is justified, and recording, for the classes where it is not, that the binding is weak and known to be weak.
Subjects and Constituencies
The subject taxonomy is five levels deep. Most enterprises operate at level four or five and reason at level two, which is why their controls and their vocabulary do not line up.
Table 3. The five levels of the subject taxonomy
| Level | What it names | Why the level matters |
|---|---|---|
| 1. Kingdom | Human, machine, organization | The coarsest cut, and the one the coverage axis uses. Binding confidence differs most sharply here |
| 2. Category | Within machine: device, workload, agent. Within human: internal and external | The level at which most policy is written, and the level at which most policy is too coarse to apply |
| 3. Constituency | Workforce, customer, partner, device, workload, agent, organization | The level at which an authoritative source either exists or does not. This is the level that predicts binding confidence |
| 4. Digital identity | The individual representation | The level at which binding is established or absent |
| 5. System objects | Accounts, identifiers, credentials | The level at which controls actually operate and entitlements are actually held |
The predictive level is the third. Whether a constituency has an authoritative source determines almost everything downstream about how well its identities can be bound. The workforce has one. Customers usually have a registration record that is authoritative for the relationship and for nothing about the person. Partners have an authoritative source inside another organization, which the enterprise cannot inspect. Workloads and agents have a creating process, which is not an authoritative source in any meaningful sense because it asserts existence rather than identity.
The four human constituencies behave differently enough to be worth separating.
- Workforce identities have the strongest available source and the weakest retention of evidence, because hiring collects substantial evidence about a person and the identity record references none of it.
- Customer identities have a source that is authoritative about the relationship and silent about the person, which is appropriate for most services and inadequate the moment a regulated obligation attaches.
- Partner identities are the most structurally difficult, because the authoritative source is inside another organization and the enterprise inherits a binding it can neither inspect nor refresh.
- Contingent workers are the constituency that most often crosses between the others, and each crossing carries forward a binding established under different conditions.
Organizations as subjects are treated thinly almost everywhere. An organization acts through identities, and the relationship between the organization and those identities is usually recorded in a contract rather than in the identity system. When a partner organization is offboarded, the identities that acted on its behalf are found by search, not by record, which is the orphaned service account failure at a larger blast radius.

Figure 2. The subject taxonomy across five levels, showing where an authoritative source exists for each constituency
Agents deserve specific treatment because they are being added to estates now. An agent is not a new kingdom. It is a workload with two properties that workloads have not previously had at scale: it acts on behalf of a subject other than itself, and it can initiate action without a human in the loop at the moment of action. The first property means an agent identity is meaningless unless the delegation chain behind it is part of the binding. The second means the usual compensating control, which is that a human is nearby and will notice, does not hold.
Watch-Out Enterprises are adopting agent identity conventions inherited from service accounts, in estates where the service account convention has already failed. In CAI’s assessment this is the single most consequential substrate decision most enterprises will make in the next three years, and most will make it by default rather than by design.
Digital Identities, Identifiers, Accounts, and the Point of Binding
This section carries the report’s principal contribution. It defines where a binding is established, why it is graded, and what the grades mean.
The point of binding
There is a moment at which an enterprise decides that a particular digital identity represents a particular subject. For a new employee it is enrollment. For a customer it is registration. For a workload it is the moment a creating process assigns an identifier. That moment is the point of binding, and it has three properties that make it the substrate’s equivalent of a control point.
- It is the only place where missing evidence can still be supplied at reasonable cost. After the point of binding, obtaining evidence about the subject means going back to the subject, which for a departed employee or a decommissioned workload is not possible at all.
- It is rarely recorded. Most estates retain the outcome of binding and discard the basis for it. The identity survives. What it rested on does not.
- It is not repeated. Bindings are established once and inherited indefinitely, through reorganizations, acquisitions, system migrations, and constituency changes, none of which re-examine the binding they carry forward.
Why binding is graded
The activity that establishes a binding is already graded everywhere it has been studied carefully. The identity assurance framework grades proofing across three levels and, in its current revision, sets out six expected outcomes and a three-step process of resolution, validation, and verification. 5 European regulation now defines compound remote onboarding procedures that combine a substantial-level electronic identification means with additional steps so that the combination meets the requirements of the high level. 6 Both treat the strength of a binding as something built up in components rather than switched on.
The discipline that has studied the underlying question longest goes further. Record linkage, the statistical problem of deciding whether two records describe the same entity, has since 1969 classified pairs into three outcomes rather than two: linked, possibly linked, and not linked, with error rates as explicit parameters of the decision. 7 The field also knows that its own optimality assumptions frequently fail in practice, and that generalizing the method beyond a small number of source files is computationally intractable, which is precisely the regime an enterprise is in when it correlates identity across dozens of systems.
In CAI’s assessment identity and access management has been answering a graded, error-calibrated question with an uncalibrated boolean for as long as it has existed, and it has done so while a rigorous treatment of the same question sat one discipline away, in plain view.
Binding confidence, S0 through S4
The scale below grades the strength of the link between a subject and the digital identity representing it. It applies to every constituency, which is its point: the standards ladder stops at the human column, and an enterprise needs one instrument that reaches across all of them.
Table 4. Binding confidence, S0 through S4
| Band | Name | What is true at this band | What it costs the controls above it |
|---|---|---|---|
| S0 | Unbound | An identity or account exists with no recorded link to any subject. The orphaned account and the unattributed service account sit here | Attribution is impossible. Certification is theater. Any entitlement decision is a decision about nobody |
| S1 | Asserted | A link exists in a record and rests on nothing beyond someone having entered it | Controls operate normally and produce results no one should rely on. This is the most dangerous band because it is invisible |
| S2 | Sourced | The link is asserted by a source trusted as authoritative for that population, with no evidence about the underlying subject | Adequate for most workforce access. Inadequate wherever the consequence of being wrong is borne outside the enterprise |
| S3 | Evidenced | The link rests on validated evidence about the subject, captured at the point of binding and retained | Supports high-consequence access and supports defensible attribution after the fact |
| S4 | Sustained | An evidenced binding re-confirmed on a defined cadence or on trigger, with the original evidence and each re-confirmation retrievable | Supports continuous evaluation and survives the events that silently break lower bands |
CAI View The band that should worry a program most is S1, not S0. An unbound account announces itself the moment anyone looks. An asserted binding looks exactly like an evidenced one in every report the enterprise produces, and it fails only when someone finally asks what it rested on.

Figure 3. The point of binding and the S0 to S4 ladder, showing what evidence each band requires and where each decays
What breaks a binding
Bindings do not expire on a schedule. They are broken by events, and the events are ordinary business occurrences that no control currently treats as substrate events.
- Departure. The subject leaves. Identities and accounts that referenced that subject as owner, creator, or accountable party now reference nobody, and only those inside a managed lifecycle are corrected.
- Constituency change. The subject remains and the relationship changes. Employee to contractor, customer to employee, partner to acquired subsidiary. The binding was established against the prior relationship and is silently repurposed for the new one.
- Acquisition and divestiture. Identities arrive with a history the acquiring enterprise cannot inspect, or depart with obligations the divesting enterprise cannot discharge. In both directions bindings cross an organizational boundary without re-examination.
- System migration. Identities are recreated in a new platform from the old platform’s records. Whatever evidence supported the original binding, if it survived at all, is rarely carried across, so a migration usually lowers binding confidence while appearing to preserve it.
- Identifier reuse. An identifier is released and reassigned. Every historical record that referenced it now points at the wrong subject, and the corruption is retroactive rather than prospective, which makes it worse.
The common property is that none of these produces an error. An enterprise that instruments these five events as substrate events, and re-examines the bindings each one touches, will improve its binding confidence faster than one that undertakes a general remediation program.
Assessing binding confidence
The assessment is tractable because it operates on classes rather than on individual identities. The method has four steps and is generally completable in weeks, not quarters.
- Enumerate classes, not identities. A class is a constituency crossed with a creation path: workforce identities created by the human resources feed, workforce identities created manually, service accounts created by the platform team, service accounts created by application teams, and so on. Most estates have between fifteen and forty classes.
- For each class, ask what the binding rested on at creation. This is a question about the creation path, not about the population, so it is answered once per class by the people who operate that path.
- Assign the band the class can evidence, not the band it intends. The gap between the two is usually the most useful output of the exercise.
- Weight by consequence, not by count. A class of two hundred privileged service accounts at S0 matters more than forty thousand customer identities at S1. Reporting by population size inverts the priority and is the most common error in the exercise.
The output is a distribution rather than a score, and it supports an investment argument directly. It names the classes where controls are producing unreliable results, and states what raising each class would require.
Accounts as objects
An account is where entitlements are actually held and where actions are actually observed, which makes it the object controls touch most and the object the substrate describes least. Treating it as a projection of an identity produces two failures directly. Orphaned accounts, meaning accounts whose identity no longer has a subject, and dormant accounts, meaning accounts whose subject no longer uses them, both become discoverable only by scanning, because the estate holds no record that would have declared them.
Table 5. Account types, and what distinguishes each
| Type | What distinguishes it | Substrate consequence |
|---|---|---|
| Standard | Held by one identity for ordinary access in a target system | The base case. Binding confidence is inherited from the identity |
| Administrative | Elevated rights, frequently a second account for the same identity | Binding to the same subject must be explicit, or privileged action becomes unattributable |
| Service | Used by software rather than by a person, often created outside any lifecycle | The accountable human is a separate fact from the account and must be recorded as one |
| Shared | Used by more than one subject by design | Binding is to a group, so attribution below the group is not available and should not be claimed |
| Ephemeral or just-in-time | Created for a bounded purpose and destroyed after | Binding must be established at creation, because there is nothing left to inspect afterward |
Substrate Failure Modes
Substrate weakness expresses itself in a small number of recurring shapes. Naming them matters because each is remediated differently, and because programs routinely apply the remedy for one to the symptoms of another.
Table 6. The six substrate failure modes, how each presents, and what each requires
| Failure mode | How it presents | What remediation requires |
|---|---|---|
| Conflation | Subject, identity, account, and credential treated as one object | A vocabulary decision and a data model change, not a tool |
| Duplication | One subject holding several identities across systems | Correlation against a stable identifier, and a decision about which identity survives |
| Orphaning | An identity or account whose subject no longer exists or is no longer known | Instrumenting departure and transfer as substrate events |
| Collision | One identifier value denoting different subjects in different scopes | Recording trust domain with every identifier, and never comparing values across domains |
| Reuse | An identifier released and reassigned, corrupting historical records retroactively | A no-reuse policy at the issuing authority, which is a discipline rather than a feature |
| Inheritance | A binding carried forward through migration, acquisition, or constituency change without re-examination | Re-establishing the binding at the crossing, or recording that it was not |
Three of these are frequently misdiagnosed as each other. Duplication and collision look alike in a report and are opposites in substance: duplication is one subject wearing several names, collision is one name worn by several subjects. Deduplication applied to a collision problem merges records that should never have been compared, which converts a correlation problem into a data integrity incident.
Inheritance is the failure mode that most rewards attention, because it is the one that accumulates. Each migration, each acquisition, and each constituency change adds a layer of bindings that were never examined, and the layers are indistinguishable from properly established bindings in every report the enterprise produces. In CAI’s assessment an enterprise that has completed two large migrations and one acquisition without treating either as a substrate event should assume its binding confidence is materially lower than its records suggest.
Watch-Out Reuse is the only failure mode that corrupts the past. Every other mode produces a wrong answer going forward. Reuse makes previously correct records wrong, which means historical attribution, retained evidence, and prior certifications all become unreliable at the moment of reassignment, with no signal that anything changed.
Resources and Operating Environments
A resource is protectable to the granularity at which it is defined and no further. This is the substrate constraint that most often surfaces as an authorization complaint. When a program reports that it cannot enforce at the granularity the business wants, the cause is usually that the resource was never defined at that granularity, and no amount of policy engine sophistication compensates.
Table 7. The four operating environments, and what each makes protectable
| Environment | What the resource is | What context exists at decision time | Granularity ceiling |
|---|---|---|---|
| Channel | A route by which access arrives: an interface, an endpoint, an API surface | Rich signal about the request, thin signal about the subject | The route, and sometimes the operation on it |
| Infrastructure | Compute, network, storage, and the control planes over them | Strong workload context, weak human context | The instance or the control plane operation |
| Data | Records, fields, objects, and the stores that hold them | Strong object context, variable subject context | The record or the field, where the store supports it |
| Application | Functions, transactions, and business objects | The richest business context of the four | The transaction, where the application exposes it |
Two consequences follow for coverage claims.
- First, an enterprise that has strong controls in the application environment and weak ones in the data environment does not have partial coverage of a single surface. It has full coverage of one surface and none of another, and the two are not interchangeable.
- Second, the same subject crossing environments carries different available context at each crossing, which is why continuous evaluation is harder in practice than its description suggests.
Resources have identifiers, and they suffer the same failures identity identifiers suffer, with less attention. A resource renamed during a migration and re-created under a new identifier severs every entitlement, certification record, and detection rule that referenced the old one. Where the old identifier is later reused for a different resource, entitlements that were never revoked become entitlements to something else. This is the reuse failure mode applied to the resource side of the substrate, and it is the least instrumented substrate defect in most estates.
Definition granularity is a decision made long before it is felt. A team that defines a resource as an application, because that is the level at which it was procured, has set the enforcement ceiling for every control that will ever be applied to it. Redefining later is not a configuration change. It means re-expressing every entitlement, every policy, and every rule that referenced the coarser object. The practical guidance is to define resources one level finer than current requirements demand, because coarsening later is cheap and refining later is not.

Figure 4. Resources across the four operating environments, with the context available and the granularity achievable in each
Watch-Out The business environment, meaning operating unit, subsidiary, and jurisdiction, is a separate coverage facet and not a fifth technical environment. Folding it into the operating environment produces coverage claims that cannot be read against regulatory scope, which is usually the reason the claim was needed.
Identity and Access Data, Verification, and Assurance
Identity and access data is the recorded form of everything above. It is the substrate object named in the Matrix as one of the two the control categories consume directly, and it is the one most often confused with the telemetry that reports on it.
What identity and access data must carry
Three properties determine whether the data can support the controls that rest on it.
- Completeness of the identifier space. Every identity and every account has at least one identifier that is stable within a stated scope, and the rules relating identifiers across scopes are recorded rather than inferred at query time.
- Retention of the basis, not only the outcome. The record states not merely that a binding exists but what it rests on and when it was established. Without this, binding confidence cannot be computed at all, and every band above S1 becomes an assertion.
- Currency proportionate to consequence. Facts that drive high-consequence decisions are refreshed on a cadence matched to how fast they change. Employment status changes faster than legal name; both are refreshed as though they were the same kind of fact in most estates.
Declarative substrate
Where identity and access data is expressed as versioned, declarative artifacts rather than as accumulated state in a directory, three things become possible that are otherwise very difficult:
- The basis of a binding can be reviewed like any other change
- Drift can be detected by comparison rather than by scanning
- The substrate for a new environment can be created deliberately instead of by accretion
This is the substrate’s share of the programmable controls argument the Root Document develops, and in CAI’s assessment it is the most credible path to raising binding confidence for machine constituencies at scale.
Two conditions have to hold for a declarative substrate to deliver those benefits.
- The artifacts must be the source rather than a description of a source, because a declaration that documents a directory instead of producing it drifts within weeks.
- The binding basis has to be expressible in the artifact, not merely the binding outcome, or the approach reproduces the same loss of evidence at higher velocity.
What a substrate assessment produces
A completed assessment produces four artifacts, and their usefulness is worth stating because programs frequently expect a score and receive something more actionable.
- A class register. Every class of identity in the estate, defined by constituency crossed with creation path, with an owner named for each. Most enterprises have never held this document and find it useful independent of anything else in the assessment.
- A band distribution. The binding confidence each class can evidence, weighted by the consequence of the access that class holds rather than by the size of the population.
- A gap statement per class. What raising a class by one band would require, expressed as an operational change, not a purchase. Many classes rise a band by retaining evidence already being collected.
- A list of unexamined inheritances. Every migration, acquisition, and constituency change in recent history whose bindings were carried forward without re-examination, which is where remediation usually has the highest return.
Identifier space quality
The identifier space is the part of identity and access data that determines whether correlation is possible at all, and its quality reduces to four properties that can be stated for any estate.
- Stability. An identifier denotes the same thing for as long as that thing exists. Identifiers derived from mutable attributes, including email addresses and usernames built from names, are not stable, and estates that use them as keys are performing correlation against a moving target.
- Scope declaration. Every identifier is recorded with the trust domain that issued it. Without this, values from different domains are compared as though they were comparable, which is the collision failure mode waiting to happen.
- Non-reuse. A retired identifier is never reassigned. This is a policy held by the issuing authority and cannot be enforced downstream by anyone consuming the identifier.
- Declared relationships. Where one identity carries identifiers in several scopes, the relationships between them are recorded, not inferred by matching at query time. Inference by matching is exactly the record linkage problem, and performing it repeatedly and implicitly is the most expensive possible way to solve it.
An estate meeting all four can correlate reliably. An estate meeting the first two and neither of the last two can correlate most of the time and cannot tell which occasions are the exceptions, which is the condition most enterprises are actually in.
Assurance, and where the ladder stops
Assurance is the property produced by verification. For human subjects, a mature graded framework exists, is regulated against, and is widely certified to. For machine, workload, agent, and organization subjects, no equivalent framework exists, because the framework that defines the levels excludes them by scope.
The consequence is not that machine identity is ungoverned. It is that machine identity is governed without a definition of what adequate would look like. Programs manage rotation, expiry, and privilege for machine identities and have no statement at all of how confidently those identities are bound to anything. The result is visible in the enterprise data: most organizations have no formally adopted policy for creating or retiring the identities their AI systems use, and a large majority reported that their organization’s definition of a privileged user applies only to human identities while a substantial share of machine identities hold privileged or sensitive access. 8
Standards Note The identity assurance framework in force limits its subject scope to natural persons. Enterprises citing identity assurance levels in policy that covers service accounts and agents are citing a framework that does not reach them. This is not a compliance risk so much as a planning risk: the absence of a target means there is no way to argue for the investment that would close it.
Substrate Coverage
The coverage axis is defined in the Matrix and elaborated here. What this report adds is the observation that binding confidence, read across the three facets, is not uniform and is predictable in its non-uniformity.
Table 8. Binding confidence typically observed, by constituency
| Constituency | Typical band | Why it sits there |
|---|---|---|
| Workforce | S2 | An authoritative source exists and is trusted for employment facts. Evidence about the person is usually collected at hire by a different function and not retained in a form the identity record can reference |
| Customer | S1 to S2 | The registration record is authoritative for the relationship and asserts little about the person. Regulated sectors reach S3 for a subset |
| Partner | S1 | The authoritative source sits inside another organization and cannot be inspected. The binding is inherited on trust and rarely re-examined |
| Device | S2 | Enrollment produces a reasonably strong binding to a managed estate, though frequently to an asset record rather than to a person |
| Workload | S0 to S1 | An identifier is assigned at creation. Nothing about a subject is established, and the accountable human is recorded outside the identity system if at all |
| Agent | S0 to S1 | Currently inheriting the workload convention. The delegation chain that would give the binding meaning is generally not part of the identity record |
| Organization | S2 | Legal entity registration provides a workable source. Rarely connected to the identities that act on the organization’s behalf |

Figure 5. Binding confidence by identity constituency, showing where the standards assurance ladder reaches and where it does not
Read across the technical environment facet, the pattern differs. The application environment usually knows most about the subject and least about the workload; the infrastructure environment the reverse. An enterprise pursuing consistent binding confidence across environments is therefore pursuing two different remediation programs, not one, and should plan them as such.
Read across the business environment facet, jurisdiction dominates. Where an enterprise operates determines what evidence it is permitted to collect and retain, which sets a hard ceiling on achievable binding confidence for some constituencies in some places. That ceiling is a legitimate constraint and should be recorded as one rather than reported as underperformance.
There is also a diagonal worth naming. The constituencies with the weakest bindings, workloads and agents, operate predominantly in the environments with the strongest technical context and the weakest subject context, which is infrastructure. The constituencies with the strongest bindings operate predominantly in the application environment, where subject context is richest. The estate is therefore best instrumented exactly where binding is strongest and worst instrumented where binding is weakest, and the two effects compound rather than offset.
For the business environment facet, two practical consequences follow. Jurisdictional limits on evidence collection set a ceiling that no investment removes, so those classes should be planned against the ceiling, not against a general target. And where an enterprise operates through subsidiaries with their own identity estates, bindings established in one subsidiary are inherited by the group without inspection, which is the inheritance failure mode operating at organizational scale.
CAI View A coverage claim expressed as a single percentage of identities verified is not a coverage claim. It averages across constituencies whose achievable bands differ by two or three levels, and the average conceals precisely the concentration that would tell a program where to act.
Common Misconceptions
- That the substrate is infrastructure. Infrastructure is bought and operated. The substrate is a condition an enterprise is in. No purchase changes it directly, though several purchases can make it measurable.
- That one identity per person is achievable and the gap is a cleanup backlog. The gap is structural, produced by acquisitions, by constituency changes, and by systems that create identities as a side effect. It is managed and measured, not closed.
- That an identity assurance level applies to a service account. The framework that defines those levels excludes non-person entities by scope. Applying the label to a service account claims a rigor that the framework does not offer and did not intend.
- That an identifier is unique because it looks unique. Uniqueness is a property of the issuing authority’s discipline. A well-formed value that is reassigned after deprovisioning is a correlation hazard wearing the costume of a primary key.
- That machine identity granularity is a vendor failure. The emerging standards deliberately permit many workload instances to share one identifier. Expecting instance-level attribution from that identifier is a misreading of the model, not a gap in it.
- That the substrate can be fixed once. Bindings are broken by ordinary business events that recur. A program that treats substrate repair as a project rather than as a standing obligation will be back in the same position within two or three years, having spent the money once.
- That more credentials mean a stronger binding. Credential strength establishes that whoever holds the credential controls the identity. It says nothing about whether the identity represents the subject it claims to. The two are independent, and conflating them lets a program report progress on the easier one.
- That verification and assurance are the same thing. Verification is an activity performed at a point in time. Assurance is the property that persists afterward, decays afterward, or was never established. Programs that measure the activity and report the property overstate what they know.
A worked example
A mid-sized insurer runs a claims platform. The example follows one human subject and one machine subject through a single year, and grades the binding at each step.
In March, an analyst joins through an acquisition. The acquired company’s directory is migrated, and her identity arrives with it. The enterprise’s authoritative source for employment now asserts that she is an employee, so her binding is S2 by the enterprise’s own account. What the enterprise does not hold is any evidence about the person: the acquired company collected it at its own hiring, retained it in a human resources system that was not migrated, and the identity record references nothing. Her binding is S2 with respect to employment and S1 with respect to the person, and no report the enterprise produces distinguishes the two.
In June, she moves to a contractor arrangement for a family reason. Her constituency changes. Her identity does not, because the change is processed as a status update on the same record. The binding that was established against an employment relationship now supports access under a different relationship, and nothing re-examined it. In S0 to S4 terms the binding did not move. In substance it should have been re-established, because the authoritative source that supported it is now authoritative for something else.
In parallel, a workload. In April an engineer creates a service account so that a reconciliation job can read the claims database nightly. The account is assigned an identifier and a credential. Its binding is S0: an identifier exists, and nothing links it to any subject at all. The engineer records his name in the change request, which is closed in July.
In September the job begins reading records it does not need. Observability detects the pattern and raises it. Attribution proceeds to the account and stops, because there is nothing beyond the account to attribute to. The change request is retrievable, but the engineer left in July, and the record shows only that he opened it. The investigation reaches the correct account within an hour and cannot say, on the evidence the substrate holds, who owns the behavior or whether it was ever intended.
In November the annual certification runs. The analyst’s manager certifies her access, and the certification is recorded as complete. It rests on a binding that was inherited from an acquisition, carried through a constituency change, and never evidenced. The service account is certified by an application owner who inherited it and who confirms that the job is still needed, which is true, and says nothing about who is accountable for it, which is the fact the certification was supposed to establish.
Two changes would have altered the year, and neither is a purchase. If the identity record created in March had carried a reference to the evidence the acquired company held, with a note of what it was and where it lived, the analyst’s binding would have been assessable at S3 rather than assumed at S2, and the June constituency change would have had something to re-examine. If the service account created in April had required an accountable identity as a field rather than a name in a change request, the September investigation would have reached a person in the same hour it reached the account. Both are record-keeping decisions taken at the point of binding, and both cost more to retrofit than they would have cost to make.
CAI View Nothing failed in this example. The governance control ran, the runtime control ran, the observability control ran, and each performed as designed. The year’s outcome was determined at two points of binding, in March and April, and at neither point did anyone record what the binding rested on.
What each audience takes away
- IAM leader. Ask for binding confidence by constituency, not a verification percentage. The distribution is the finding, and the weakest constituency sets what your other investments can achieve.
- Control architect. The substrate belongs on your architecture as an explicit layer with named owners, not as an assumption beneath it.
- Control engineer. Retain the basis of a binding, not only its outcome. Nothing above S1 is computable without it.
- Solution designer. Define resources at the granularity you intend to enforce at. Enforcement cannot exceed definition.
- Market analyst. Vendors that expose binding confidence as a queryable property of the identity record are doing something structurally different from vendors that report proofing workflow status. The distinction will separate the category.
- Academic researcher. The record linkage literature addresses this problem rigorously and has not been applied to enterprise identity. The transfer is open work.
- Consultant. The S0 to S4 grading is assessable in weeks from existing data and produces a defensible investment case where a maturity score does not.
Conclusion
The substrate is the set of facts every identity and access control assumes before it runs, and the strength of those facts is a graded property that almost no enterprise measures. The link between a subject and the identity representing it is the substrate’s unit of control. It is established at a point, it is rarely recorded, it is never repeated, and every governance decision, runtime decision, and attribution inherits whatever it was worth at the moment it was made.
The graded instrument this report contributes, S0 through S4, exists because the standards ladder that grades human proofing stops at the boundary of the human constituency, and the majority of a modern estate lies on the other side of that boundary. An enterprise that grades its binding confidence by constituency learns two things quickly: where its controls are producing results that should not be relied upon, and which of its planned investments cannot succeed until the substrate beneath them is repaired.
None of this argues for delay. The controls an enterprise operates today are the controls it has, and they do useful work at every band. What the grading changes is the confidence with which their results are used and the order in which the next investments are made. A program that knows two hundred privileged service accounts sit at the unbound band has a specific, small, and fundable problem. The same program without the grading has a general sense that machine identity needs attention, which has funded very little anywhere.
The three control categories treated in the companion Foundation Reports are each strong disciplines. None of them can be more correct than this report’s subject matter allows. In CAI’s assessment the programs that will separate themselves over the next several years are not the ones that buy the most capable controls. They are the ones that find out what their controls have been resolving against.
Acronym Key and Glossary
| CSP | Credential service provider. |
|---|---|
| IAL | Identity assurance level. |
| AAL | Authenticator assurance level. |
| FAL | Federation assurance level. |
| DID | Decentralized identifier. |
| VC | Verifiable credential. |
| SCIM | System for Cross-domain Identity Management. |
| SPIFFE | Secure Production Identity Framework for Everyone. |
| WIMSE | Workload Identity in Multi System Environments, an IETF working group. |
| PID | Person identification data, as used in European digital identity wallet specifications. |
| Binding | The established link between a subject and the digital identity representing it. |
| Binding confidence | The graded strength of that link, S0 through S4. |
| Point of binding | The moment at which an enterprise decides that a digital identity represents a particular subject. |
| Identifier space | The set of identifiers across an estate together with the rules relating them. |
| Constituency | The class of subject an identity represents. Full definitions appear in the Definitions section. |
Evidence
Standards and specifications
- NIST SP 800-63-4 and SP 800-63A-4, Digital Identity Guidelines, final July 2025, superseding SP 800-63-3.
- ISO/IEC 24760-1:2025, 24760-2:2025, and 24760-3:2025, A framework for identity management.
- W3C Verifiable Credentials Data Model 2.0, Recommendation, 15 May 2025.
- W3C Decentralized Identifiers 1.0, Recommendation, 2022, and Decentralized Identifiers 1.1, Candidate Recommendation Snapshot, 5 March 2026.
- IETF draft-ietf-wimse-identifier-03 and draft-ietf-wimse-arch-08.
- Regulation (EU) 2024/1183 and Commission Implementing Regulation (EU) 2026/798.
Academic
- Fellegi and Sunter, A Theory for Record Linkage, Journal of the American Statistical Association, 1969.
- Christen, Data Matching, Springer, 2012.
- Getoor and Machanavajjhala, Entity Resolution: Theory, Practice and Open Challenges, VLDB, 2012.
- Binette and Steorts, (Almost) All of Entity Resolution, Science Advances, 2022.
Enterprise and practitioner
- Cloud Security Alliance and Oasis Security, The State of Non-Human Identity and AI Security, January 2026.
- CyberArk Identity Security Landscape, 2025.
- Rubrik Zero Labs, non-human identity ratio research, 2026.
- Palo Alto Networks, 2026 Identity Security Landscape.
- OWASP Non-Human Identities Top 10, 2025.
Method note
Claims in this report were verified against primary sources. Where a claim could not be traced from a secondary source to a primary one, it was removed rather than qualified. No single machine-to-human identity ratio is asserted anywhere in this report.
Related Reading
- IAM-RD-01, the Root Document, for the Control and Coverage Matrix in full.
- IAM-F-GOV-01 for the lifecycle that creates and retires the objects defined here.
- IAM-F-RTM-01 for the credential model and the runtime access decision.
- IAM-F-OBS-01 for attribution and the operational data foundation.
- IAM-CAP-IDP-01 for (upcoming) identity proofing and verification methods.
Publication Notes
This report defines several terms that earlier Foundation Reports carried as working definitions pending its publication. Where the definitions here differ from those working statements, the definitions here govern. The published reports stand as issued, and the differences are recorded in the Companion Guide.
Three terms are CAI coinages introduced by this report and are not drawn from any external standard: identity binding as the substrate’s unit of control, binding confidence as its graded expression, and the point of binding as the moment at which the link is established. The S0 through S4 band names are likewise CAI’s.
The report deliberately avoids naming a single figure for the ratio of machine to human identities. The reasons are given in the third observation and in the associated note.
Notes
- NIST SP 800-63A-4, Introduction: within the scope of the guidelines the terms subject and person refer to natural persons, and not to non-person entities, organizations, or things. NIST SP 800-63-4 was published final in July 2025 and supersedes SP 800-63-3.
- Published machine-to-human ratios in current circulation span roughly an order of magnitude, Rubrik Zero Labs reports 45 to 1 for the modern enterprise, CyberArk’s 2025 Identity Security Landscape 82 to 1, Palo Alto Networks’ 2026 Identity Security Landscape 109 to 1, and Entro Labs 144 to 1 for cloud-native and DevOps environments in the first half of 2025. The Cloud Security Alliance has itself cited figures at both ends of that range in publications roughly two months apart. Survey-derived and telemetry-derived figures are not equivalent, and population scope differs materially between them.
- ISO/IEC 24760-1:2025, Core concepts and terminology, third edition, which cancels and replaces the 2019 second edition. Parts 2 and 3 were likewise reissued in 2025. Part 1 treats the distinction between the terms identity and identifier as a core concept, and ISO describes the standard as clarifying the distinctions between identity, identifier, attributes, and other core concepts.
- IETF draft-ietf-wimse-identifier-03, July 2026, provides that an identifier assigned to a workload should not be reassigned to a different workload unless the trust domain’s policies explicitly intend it, and that multiple workload instances may share the same Workload Identifier. The same draft notes that a workload identifier has meaning only within the scope of a specific issuer.
- NIST SP 800-63A-4 names identity resolution, evidence validation, attribute validation, identity verification, identity enrollment, and fraud mitigation as the expected outcomes of identity proofing, and structures proofing as three steps: resolution, validation, and verification.
- Commission Implementing Regulation (EU) 2026/798 of 7 April 2026, laying down rules for the application of Regulation (EU) No 910/2014 as regards reference standards and specifications for the remote onboarding of users to European Digital Identity Wallets by electronic identification means conforming to assurance level substantial in conjunction with additional remote onboarding procedures where the combination meets the requirements of assurance level high.
- I. P. Fellegi and A. B. Sunter, “A Theory for Record Linkage,” Journal of the American Statistical Association 64(328), 1969, pages 1183 to 1210. The framework classifies record pairs by likelihood ratio into linked, possibly linked, and non-linked, and is optimal under stated assumptions. For a modern treatment including the limits of the conditional independence assumption, see Binette and Steorts, “(Almost) All of Entity Resolution,” Science Advances, 2022.
- Cloud Security Alliance and Oasis Security, The State of Non-Human Identity and AI Security, January 2026: 78 percent of organizations lack formally adopted policies for creating or removing AI identities, and 79 percent of respondents report feeling ill-equipped to prevent attacks through non-human identities. CyberArk 2025 Identity Security Landscape, conducted by Vanson Bourne across 2,600 cybersecurity decision makers in 20 countries: 88 percent reported that the definition of a privileged user in their organization applies solely to human identities, while 42 percent of machine identities hold privileged or sensitive access.
Disclaimer
© 2026 Control Architecture Institute Inc. All rights reserved. Control Architecture Institute Inc. (Control Architecture Institute or CAI) is a research institute. The Library, of which this report is a part, consists of the opinions of CAI’s research team, which should not be construed as statements of fact. While the information contained in this publication has been obtained from sources believed to be reliable, CAI disclaims all warranties as to the accuracy, completeness, or adequacy of such information. Although CAI research may address legal and financial issues, CAI does not provide legal or financial advice and its research should not be construed or used as such. Your access and use of this publication are governed by CAI’s terms of use. CAI prides itself on its independence and objectivity. Its research is produced independently by its research team without input or influence from any third party. This publication may not be reproduced or distributed in any form without CAI’s prior written permission. CAI research may not be used as input into, or for the training or development of, generative artificial intelligence, machine learning, algorithms, software, or related technologies.