Containing a compromised agent
When a user's delegated agent is hijacked, spoofed, or tricked into acting beyond its scope, containing it: cutting off the agent's standing authority before the harm compounds.
The impact of agents
As agents hold standing authority and act over time and across services, a compromised or spoofed agent becomes able to act at a distance. It acts without notice, at machine speed. The takeover of an internet-facing agent, the impersonation of a trusted one, and an over-privileged agent tricked into exceeding its scope are all recognized attack classes, and an agent that holds a user's authority is a standing target for them. A compromise that happens once can keep acting for as long as the credential remains valid, because nothing at the point of action tells the service it is no longer dealing with the agent the user authorized.
What must be verified
An agency needs two things when an agent holds standing authority to act for a user: confidence that the agent presenting that authority is the one the user authorized and not a look-alike, and the ability to withdraw a compromised agent's authority in the window that matters, not whenever the credential would have lapsed on its own. The first is a binding between the delegation and a verifiable agent identity; the second is revocation that takes effect within that window. The identity issuer must hold that binding, and the authorization server or agency must operate the propagating revocation.
Protecting access
Containment that strands the user is its own exclusion. Someone who depends on an agent because they cannot operate a service unaided is denied the outcome if a suspected compromise freezes their authority without notice or a route back. A person already wary of the system reads that lockout as the system turning on them.
Keeping the path open
- Pair containment with an assisted re-authorization route the user can complete through a supported channel.
- Keep a first-class non-agent fallback, so a contained agent never means no service.
- Degrade gracefully: treat a suspected compromise as grounds for stepped-up confirmation, and keep the underlying entitlement available while confirming.
- Make the containment notice and the re-authorization route operable by keyboard and assistive technology, stating what is paused and the way back in plain language.
Response surface
Revocation reaches every relying service within a second, while the user sees a plain-language pause with a route back rather than a silent lockout.
Its recent activity was unusual, so we reduced its access while we check. Your entitlements, lodgment, refunds, and deadlines are unaffected.
A freeze with no explanation reads, to someone already wary of the system, as the system acting against them. Containment names what happened, leaves the entitlement intact, and keeps a route open.
Revocation confirmed by 3 of 3 relying services.
Tie the authority to a verifiable agent identity, and a spoofed agent fails the moment it presents itself. Revocation spreads in near real time, pulling a compromised agent before its next action. Neither case strands the user who relied on it.
Maturity
- Emerging
For the revocation layer, where token revocation, a continuous-evaluation profile, and a privacy-preserving credential status list are ratified standards, deployable today and not yet routine practice for agent authority.
- Frontier Headline
For verifying that the acting agent is the one the user authorized, where interoperable agent identity is proposed but not yet built.
Precedents
OAuth token revocation (RFC 7009), OpenID CAEP, and W3C Bitstring Status List. RFC 7009 lets a client tell the authorization server that a credential is no longer valid, and on revoking a refresh token the server should invalidate the tokens issued under the same grant. That does not expire an already-issued short-lived token at the services relying on it, so a continuous-evaluation profile pushes the revocation event to those services, and a status list marks a verifiable credential revoked or suspended. All three are ratified standards.
OWASP LLM06 Excessive Agency, and MITRE ATLAS. OWASP describes the damage an over-privileged agent can do when its output is manipulated, and prescribes least privilege, execution in the individual user's context, and authorization enforced by the downstream system instead of the agent. MITRE ATLAS catalogs credential abuse against AI systems among its documented adversary techniques. The controls limit how far a compromised agent reaches before it is revoked, and the technique catalog is drawn from observed practice.
Five Eyes joint guidance on agentic AI services. The guidance names the confused-deputy risk, in which a trusted, over-privileged agent is misused to perform unauthorized actions, and recommends trigger-action protocols that restrict an agent's permissions when unexpected behavior emerges, a prohibition on agents altering their own privileges, and trust scoring that drops on anomalous behavior. Six national cyber authorities describe the containment controls, and the risks they set out are illustrative scenarios rather than documented incidents.
What carries over to agent use
Containment separates cleanly into a layer that is deployable and a layer that is not. Deployable now: bind an agent's authority to a credential that can be revoked, and wire revocation to propagate to relying services in near real time, so pulling a compromised agent's authority is a matter of seconds rather than the life of a token; hold the agent to least privilege and downstream-enforced authorization so a single compromise cannot reach every service at once. Not yet deployable: a verifiable binding that lets a relying service confirm the acting agent is the one the user authorized, which depends on agent-identity standards still in draft.
No documented incident yet shows a government-service user agent being hijacked to transact; the observed cases sit in enterprise and developer-tool contexts. The honest framing is that the threat is projected from adjacent domains rather than seen in this one, that the containment primitives are real and worth building ahead of the incident, and that the identity layer that would complete them is the part the field has not yet solved.
Where things go wrong
The failure mode is a standing delegation that keeps acting after the agent behind it has been taken over, impersonated, or tricked into exceeding its scope. An agent holding authority across several services can be driven to act on all of them before any single action looks wrong, because each action is locally valid while the agent is globally compromised. The user often discovers the harm only once it has compounded. A spoofed agent fails at presentation once its authority is bound to a verifiable agent identity, and revocation that propagates in near real time pulls a compromised agent's authority before the next action rather than after the credential eventually expires. Containment can itself be turned into the attack: a forged compromise signal, or an institution's own over-caution, cuts a person off from the authority they depend on. That is why a containment event carries a plain-language notice, an assisted route back, and a record the user can contest.
Sources
7 references
The instrument, the operating deployment, or the official record itself.
- RFC 7009 — OAuth 2.0 Token Revocation
- OpenID CAEP 1.0 — Continuous Access Evaluation Profile
- W3C Bitstring Status List v1.0 (credential revocation)
- OWASP LLM06:2025 — Excessive Agency
- Careful adoption of agentic AI services — Five Eyes joint cybersecurity guidance (ASD ACSC, CISA, NSA, CCCS, NCSC-NZ, NCSC-UK)
Writing about the subject rather than the framework itself, including vendor commentary.