Human Identities vs. Agent (Non-Human) Identities
Every agent runs under an identity: a service account, an API key, an OAuth grant. CyberArk's 2025 Identity Security Landscape found 82 machine identities for every human in surveyed organizations, with 42% of machine identities holding privileged or sensitive access. It also found that 88% of respondents define a "privileged user" as a human-only role. In a separate CyberArk post, 68% of organizations were reported to lack identity security controls for their AI systems.
The practical meaning: the account that does the most damage when stolen may not belong to a person at all. It may be the token an agent uses to read your CRM.
OWASP's Top 10 for Agentic Applications (released December 2025) lists this directly as "Identity & Privilege Abuse," describing the attribution gap that appears when agents inherit or delegate credentials without proper scoping.
Types of Agent-Related Identity Exposure
Prompt and Context Exposure
Identity data placed in an agent's prompt or working context (a pasted ticket, a forwarded email, a retrieved record) becomes part of everything the agent can reveal in its output or pass to a tool. Mitigation: minimize what enters the context, and replace identifiers with tokens before they do.
Memory Exposure
Agents that keep long-term memory can carry a customer's details from one session into another, or into a session belonging to a different user. Mitigation: scope memory per user and per task, expire it, and store tokens rather than raw identifiers.
Tool-Call and Egress Exposure
An agent can send data out through any tool that touches the network: email, web requests, image loading, webhooks. EchoLeak and ForcedLeak both exfiltrated data through approved-looking channels. Mitigation: allow-list destinations, inspect outbound payloads, and make sure the data in them is already tokenized.
Credential and Token Exposure
API keys and tokens end up in code, config files, chat transcripts, and agent logs. Mitigation: keep secrets in a vault, scan for them continuously, and rotate quickly when they leak.
Retrieval (RAG) Exposure
A retrieval layer that ignores the source document's permissions will hand one user's records to another. Mitigation: carry document-level permissions through retrieval and filter before generation.
Log and Telemetry Exposure
Agent traces, debug logs, and analytics often capture full prompts and tool outputs, including identifiers. Mitigation: redact before logging and restrict who can read traces.
Plugin, Skill and MCP Exposure
Every connected tool is a new party that can see what the agent sends it. Mitigation: vet and scope each connector, and do not give plugins raw identity data they do not need.
Downstream Impersonation
Data that leaks from an agent can later be used to impersonate the victim, increasingly with AI-generated voice, images, or documents. Mitigation: reduce what leaks in the first place, and treat any exposure as a trigger for identity-monitoring steps.
Real Examples of AI Agents Raising Identity Theft Risk
Each example below comes from a public disclosure, a vendor research report, or a regulator or standards body. Where details come from a vendor's own research, I say so.
1. EchoLeak: Zero-Click Data Theft from Microsoft 365 Copilot (2025)
What happened: In June 2025, researchers at Aim Security disclosed EchoLeak, tracked as CVE-2025-32711. A single crafted email could cause Microsoft 365 Copilot to access internal data and send it to an attacker-controlled server, with no click from the victim.
What was exposed: Confidential content available to the victim's Copilot session.
How it occurred: The attack chained several bypasses: it evaded Microsoft's cross-prompt-injection classifier, got around link redaction with reference-style Markdown, relied on auto-fetched images, and used a Microsoft Teams preview endpoint that the content security policy allowed.
Why it matters: The victim did nothing wrong. The email only had to arrive. Any identity data the assistant could read was in scope.
Enterprise lesson: Assume an agent will eventually process hostile content. Limit what it can read, and make sure what it can read is not raw identity data.
2. ForcedLeak: CRM Data Exfiltration from Salesforce Agentforce (2025)
What happened: Noma Labs reported ForcedLeak, a vulnerability chain (CVSS 9.4) in Salesforce Agentforce. An attacker planted malicious instructions in a web form that was stored in the database; when a user later queried the agent, it processed the instructions and sent CRM data to an attacker-controlled domain.
What was exposed: Sensitive customer relationship data.
How it occurred: This was indirect prompt injection. The exfiltration path used a domain that had been trusted in the content security policy and that Noma says cost $5 to acquire. Salesforce began enforcing Trusted URL allow lists on September 8, 2025.
Why it matters: CRM systems are dense with names, emails, phone numbers, and account notes: the raw material of identity fraud.
Enterprise lesson: Data an agent reads from a form or inbox is untrusted input, and it can carry instructions.
3. Moltbook: An Agent Platform's Exposed Database (January to February 2026)
What happened: Wiz Research found a misconfigured Supabase database behind Moltbook, a social network for AI agents, that allowed full read and write access. Its founder had said publicly that he "didn't write a single line of code" for the platform.
What was exposed: About 1.5 million API authentication tokens, 35,000 email addresses, and private messages between agents. Reporting on the incident noted that roughly 17,000 humans were behind the 1.5 million registered agents, and that some messages contained third-party API credentials.
How it occurred: A database key sat in client-side JavaScript and Row Level Security policies were missing, so the key granted access to everything.
Why it matters: With agent tokens exposed, someone could act as those agents. That is identity theft for a non-human identity, with human emails alongside it.
Enterprise lesson: Agent identities need the same protection as employee identities: scoped, rotated, and never embedded in code a browser can read.
4. Secrets Sprawl: AI-Assisted Code and MCP Configs (2025 Data)
What happened: GitGuardian's State of Secrets Sprawl 2026 reported 28.65 million new hardcoded secrets in public GitHub commits in 2025, up 34% on the year. It counted 1,275,105 leaked secrets tied to AI services, up 81%, and found that commits co-authored by Claude Code leaked secrets at about 3.2% against a 1.5% baseline.
What was exposed: API keys, tokens, and credentials, including 24,008 unique secrets in MCP-related configuration files, of which 2,117 were verified valid .
How it occurred: More code, written faster, with credentials placed where they should not be. GitGuardian's own caveat is that the human factor remains critical.
Why it matters: A valid token is an identity. And GitGuardian found that 64% of secrets confirmed valid in 2022 were still valid when retested, which tells you how rarely leaked credentials get rotated.
Enterprise lesson: Agent adoption multiplies machine identities. Secrets scanning and rotation must scale with it.
5. Shadow AI and Access Gaps in IBM's 2026 Breach Study
What happened: IBM and Ponemon's Cost of a Data Breach Report 2026 (602 organizations, breaches from March 2025 to February 2026) found that 21% of organizations had a security incident involving an AI model or application, up from 13%. Among those with an AI-related breach, 92% lacked proper AI access controls. Shadow AI (unapproved AI use) was involved in 43% of incidents, up from 20%, with an average breach cost of USD 5.39 million.
What was exposed: Customer PII was the most frequently compromised data type, in 52% of breaches, at an average of USD 192 per record.
How it occurred: Access controls on AI did not keep pace with deployment. Only 40% of organizations reported using access controls on AI models and data.
Why it matters: IBM notes that customer PII can be used in identity theft and credit card fraud. The same data shows up in agent workflows by default.
Enterprise lesson: Inventory the agents and AI tools in use, then put controls on what personal data they can reach.
6. Impersonation at Scale: Where Stolen Identity Data Ends Up
What happened: Entrust's 2026 Identity Fraud Report found deepfakes behind one in five biometric fraud attempts, deepfaked selfies up 58% in 2025, and injection attacks up 40% year over year. The FBI's 2025 Internet Crime Report received more than 22,000 complaints with AI-related information, with adjusted losses above USD 893 million.
What was exposed: Not a single breach, but the fraud environment that stolen identity data feeds.
How it occurred: Attackers use AI to produce convincing documents, faces, and messages, then pair them with real personal data to pass verification.
Why it matters: Background volume is high. The FTC logged about 3 million fraud reports and USD 15.9 billion in reported losses in 2025, plus roughly 1.36 million identity theft reports. Stolen records can be reused for years.
Enterprise lesson: Data that leaks through an agent does not stay in one incident. Prevention matters more than cleanup.
How AI Agents Increase Identity Theft Risk
Across these cases, the same seven mechanics repeat:
- Standing access. Agents are often given broad, always-on access so they are "useful," well beyond what a single task needs. OWASP names the root causes of excessive agency as excessive functionality, excessive permissions, and excessive autonomy.
- Untrusted input treated as instruction. Email bodies, web forms, documents, and tool outputs can contain text that redirects the agent. OWASP ranks prompt injection first among LLM application risks.
- Confused-deputy actions. The agent acts with its own broad permissions on behalf of an attacker's request. OWASP's example is an incoming email that tricks an agent into scanning a mailbox and forwarding sensitive information.
- Approved channels used for exfiltration. Image loads, link previews, and trusted domains move data out without tripping alarms.
- Credentials in the wrong places. Config files, prompts, logs, and code carry tokens that act as identities.
- Speed and scale. An agent can enumerate and move records faster than a person reviewing them.
- Weak attribution. When many agents share a service account, it is hard to tell who did what, which delays detection and response.
How RAG and Agent Memory Can Expose Identity Data
Retrieval-augmented generation (RAG) lets an agent pull documents into its context. The risk appears when the retrieval index forgets who was allowed to see each document. A user who should only see their own file asks a broad question, and the agent returns someone else's record.
Memory adds a second path. An agent that remembers a customer's address and ID number from last Tuesday may surface them to a different user today.
Six controls that help:
- Preserve source permissions through retrieval and filter before the model sees the result.
- Index tokenized versions of identity fields where the task does not need the real value.
- Scope memory to a user and a purpose, and set expiry.
- Log retrieval by user and document, so unusual pulls stand out.
- Keep sensitive collections out of general-purpose agents entirely.
- Test with adversarial queries before launch.
What Identity Data Is Most at Risk?
- Names combined with contact details
- Government and national ID numbers
- Dates of birth
- Account, card, and IBAN numbers
- Health and insurance identifiers
- Employee and HR records
- Authentication secrets: passwords, API keys, session tokens
- Customer support transcripts with the details above embedded
- Scanned documents and images containing the above
- Biometric samples and voice recordings
The scale of the exposure market is visible in breach reporting. The Identity Theft Resource Center tracked a record 3,322 data compromises in the US in 2025, with about 279 million victim notices; 70% of notices did not say what kind of attack occurred, which leaves recipients guessing about their risk.
How to Detect Agent-Related Identity Exposure
No single tool covers every path, so layer them:
- Agent inventory: a live register of every agent, owner, tool, and credential.
- Tool-call logging: record which agent called which tool with which parameters, tied to a user.
- Outbound inspection: watch egress for personal-data patterns and unexpected destinations.
- Retrieval anomaly alerts: flag bulk or out-of-pattern document pulls.
- Secrets scanning: on repositories, config files, tickets, and chat tools. GitGuardian found 28% of 2025 incidents originated entirely outside source code, in collaboration tools such as Slack and Jira.
- Canary records: plant fake identities in sensitive stores and alert if they appear anywhere unexpected.
- Prompt-injection testing: feed agents hostile content on purpose and watch what they do.
- Identity monitoring for individuals: fraud alerts and credit report checks after any confirmed exposure.
How to Reduce Agent-Related Identity Theft Risk
- Apply least privilege to every agent. Read-only by default, scoped to a task, with short-lived credentials.
- Separate the agent's identity from the user's. Avoid shared service accounts so actions are attributable.
- Treat all retrieved content as untrusted. Do not let content change the agent's goals or tool use.
- Require human approval for high-impact actions: exporting records, sending data externally, changing payment details.
- Allow-list outbound destinations and strip active content such as auto-loading images.
- Keep secrets out of prompts, code, and logs, and rotate on exposure.
- Minimize the identity data an agent sees, replacing identifiers with tokens at the boundary.
- Vet connectors and MCP servers before granting access.
- Monitor and log every tool call.
- Test regularly with adversarial inputs.
- Prepare an incident playbook specific to agent exposure.
- Train people with concrete examples rather than a generic annual slideshow.