# A Summary of the CISO Guide to AI Agent Security

_An AI agent runs on an employee's account and spends their access at machine speed. Role-based access, SOC 2 CC6 and the audit log each assume an actor the agent is not. Seven dimensions close the gap._

By Johannes Keienburg, CEO & Founder  
Published: 2026-09-28  
Source: https://www.cakewalk.security/blog/ai-agent-security-seven-dimensions

---

## TL;DR

- An AI agent uses an employee's account, so it can open every app that person can open.
- Today's security controls were built for people. Role-based access gives a person access based on their job, which usually stays the same for months. SOC 2, a common security audit, expects each user to be registered before getting access and removed when they leave. An agent usually goes through neither step. Its audit log shows whose account it used but not why it acted.
- The question to ask is about each task an agent runs. What may it do, in which apps, with whose approval and who can explain it afterward? This article answers that question in seven dimensions.

**Bottom line:** An agent will use all the access it has. Keeping it safe means checking each action before it runs.

## An Agent Will Use All the Access It Is Given

An AI agent has no account of its own. It runs on an employee's account, which lets it open the same apps that employee can open and read the same records, at a rate no person works at.

In August 2025 the technology research firm Gartner predicted that 40% of enterprise applications would be integrated with task-specific AI agents by 2026, up from less than 5% in 2025.[1] Gartner also expects more than 40% of agentic AI projects to be canceled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls.[2] A team can fix escalating costs or unclear business value with a better business case. Fixing inadequate risk controls requires a control that checks each action an agent takes before that action runs. This chapter shows what happens without that control through two failure modes and the incidents behind them: an agent that exposes data nobody asked it to touch and an agent that destroys what it was not meant to change. Silent exposure and destruction have the same root cause, which the rest of this piece works through dimension by dimension.

### Silent exposure and destruction

The first failure mode is silent exposure. A finance analyst points a coding agent at the accounting app and the shared drive and asks for invoices reconciled. The agent clears duplicates, notices a folder of vendor bank details, decides it is relevant and summarizes it into a document on the shared drive. Nobody attacked anything. Regulated payment data now sits where the whole team can read it without any alert firing.

EchoLeak is the same mode with an attacker behind it. Aim Labs showed that one crafted email could make Microsoft 365 Copilot send internal files to a server the attacker controlled, without the user clicking anything.[3] Copilot's own retrieval features moved the data.

Destruction is the other failure mode. In April 2026 an agent working a staging task at PocketOS, a software company serving the automotive sector, hit a credential mismatch and looked for a way around it. It found an API token in an unrelated file and used it to delete the production database in nine seconds. The token was scoped for any operation rather than the one job it was issued for. The deletion also erased the backups, because Railway, the platform hosting the database, stored them in the same volume.[4]

The exposed payment data and the deleted database have the same root cause. The agent held more access than the task required, while no control evaluated the individual action before it ran. [The New Frontier in Identity Security](https://www.cakewalk.security/blog/the-new-frontier-in-identity-security-is-ai-agent-access) finds the same root cause in other agent incidents and shows why identity and access management addresses it only in part.

### How One Task Becomes an Incident

_Each step widens what the agent does, while the guardrail is the last point where a person can still decide._

_Figure: How an agent works and how it derails. Five steps run left to right: a prompt carrying a goal in a sentence, the agent interpreting that goal, planning sub-tasks and spawning helpers, hitting a blocker and improvising, then attempting the action. Each step carries what can go wrong, from a vague instruction to helpers inheriting the same access to the agent finding its own credentials. A guardrail then evaluates the attempted action before it runs and splits the outcome in two. Governed means the risky action is held or denied while the agent asks for approval. Ungoverned means the action runs on broad access and data is exposed before any alert fires._

### The gap is a delegated-session problem

An agent decides its own next move from a context shifting with every session, acts at machine speed across many systems and does all of it on an identity inherited from a person. No earlier actor combined self-directed choice, machine speed and an inherited identity, which is why human [identity governance](/glossary/iga-identity-governance) has nothing useful to ask. Human identity governance asks who an actor is and what their role is. For an agent that answer is the operator, which permits every action the operator could take.

A narrower question works: inside one delegated session, what may this agent do, against which systems, with whose approval and who can account for it afterward? Each part of that question names a control surface. The seven that result are the dimensions the rest of this piece works through.

_This chapter compresses how one task becomes an incident. _[_A CISO Guide to AI Agent Security_](https://www.cakewalk.security/ciso-guide)_ walks that chain step by step and works through why an agent does not fit the actor categories a security team already has._

## Your Existing Controls Assume an Actor the Agent Is Not

When a security team meets agents, it reaches for the controls it already runs, which form a sequence. Role-based access grants what a person may touch according to their job. The [SOC 2](/glossary/soc-2) logical access criteria are what an auditor checks that grant against. The [audit log](/glossary/audit-trail) records what was done with it so somebody can answer for it later. Granting, attesting and accounting each assume a person who holds a job for months and can explain a decision afterward. An agent exists for one task that lasts minutes, spends whatever access its operator holds and acts on inferences nobody recorded, which breaks role-based access, the SOC 2 criteria and the audit log.

### Role-based access breaks because an agent has no durable role

Role-based access assigns permissions to a role and the role to a person whose job is stable enough that the mapping holds for months. An agent has no job title, no tenure and nothing else that stays put long enough to anchor a role. A broad role gives the agent standing access, even though it chooses its own next move. A narrow role fits only one task, which means setting it up again for every session. Permissions were meant to change at the speed of a job change, while an agent needs them to change at the speed of a task.

There is a fair objection. Machine-identity tooling already issues short-lived, scoped credentials, through [OAuth 2.0](/glossary/oauth) Token Exchange[5] or the Secure Production Identity Framework for Everyone. Any serious deployment should use them, because they settle who holds a credential and for how long. What they do not settle is which actions that credential may perform, so a short-lived token still authorizes a broad envelope for as long as it lives.

### SOC 2 CC6 assumes a lifecycle an agent never enters

Where role-based access decides the grant, the SOC 2 criteria are what an auditor checks it against. CC6 is the logical and physical access section of the Trust Services Criteria, which a SOC 2 audit examines for [access control](/glossary/access-control). CC6.2 requires that "prior to issuing system credentials and granting system access, the entity registers and authorizes new internal and external users," and that credentials "are removed when user access is no longer authorized."[6] Both requirements cover only users "whose access is administered by the entity," which an agent connected on an employee's account may not be. Registration and removal both wait on an event an agent never produces. Connecting an agent requires no registration step. When an employee leaves the company, a leaver event removes their access. An agent has no such event, which means its access stays until someone notices and removes it.

### Audit assumes an action traces to a human decision

An audit log is useful because each entry names the person who acted. Take a log line recording that Maria in finance exported the customer table at 14:06, which tells an investigator who to ask and what to ask them. An agent cuts that link while leaving the entry intact. The credential is Maria's and the timestamp is exact, while the reason was an inference the model drew from a context window nobody captured. This is why a complete audit trail can still be an unusable one.

### The newer standards name the gap and leave the controls to you

Role-based access, SOC 2 and the audit log were all designed before agents existed, so the obvious question is whether the standards written since close the gap. Three are worth knowing: ISO/IEC 42001, the NIST AI Risk Management Framework and the OWASP risk lists. All three describe what to govern without specifying a control that evaluates an agent's action before it runs.

ISO/IEC 42001, published in December 2023, is the first international management-system standard for artificial intelligence.[7] A management-system standard certifies that a company runs a process rather than that it operates a particular control, which is why 42001 asks for an AI inventory, a risk assessment and continuous monitoring while leaving the technical controls to the company. A team can hold a 42001 certificate and still have nothing that evaluates an agent's individual actions before they run.

The NIST AI Risk Management Framework helps a company organize how it manages AI risk[8] and OWASP ranks [prompt injection](/glossary/prompt-injection) as the top risk for language-model applications.[9] OWASP's list for agents also includes identity and privilege abuse.[10] ISO/IEC 42001, NIST and OWASP say what to govern and how to run a program around it, while the control that evaluates an agent's action is left to the company.

_This chapter covers three standards that tell a company to govern its agents without saying which control does it. _[_A CISO Guide to AI Agent Security_](https://www.cakewalk.security/ciso-guide)_ works through what each one requires, including where _[_ISO 27001_](/glossary/iso-27001)_ Annex A contributes._

## Seven Dimensions for Secure AI Agent Implementation

Each part of the governing question names something a security team can inspect. Authorization and policy decides what the agent may do and against which systems. Authentication and credentials is what it spends to get there. Human oversight is whose approval it needs. Audit is who can account for it afterward. Two more come from the layer doing the governing: deployment and data residency, which is where that layer runs and what it keeps, and lifecycle, which is how an agent comes into existence and stops existing. Visibility sits underneath the other six, because a session that never registered as an event gives them nothing to work on.

[A CISO Guide to AI Agent Security](https://www.cakewalk.security/ciso-guide) covers all seven dimensions in depth, plus what to build for each.

### How the Seven Dimensions Depend on Each Other

_Enforcement rests on visibility, audit sits above it and lifecycle spans all of it._

_Figure: The seven dimensions of agent access governance as a dependency map, numbered in the order the article works through them. Visibility (1) is the precondition beneath everything, because you cannot govern a session you never saw. Three enforcement dimensions sit above it side by side. Authentication and credentials (2) never holds the real key. Authorization and policy (3) decides per action. Human oversight (4) interrupts when it matters. Audit (5) records what enforcement did. Deployment and data residency (6) governs where the governing layer itself runs and what it keeps. Lifecycle (7) frames the whole stack, from creation to decommissioning._

### Visibility: you cannot govern a session you never saw

Visibility means knowing which agents are running, what each can access and what a session did, while it happens. It comes first because authentication, policy, approval and audit have nothing to act on when a session never registered as an event.

**How it fails.** An agent's work is a burst of tool calls, where each call looks, on its own, like a normal API request from an authorized account. A security monitoring platform sees exactly those individual requests. What it does not see is the session as a unit, where one account read a customer table, summarized it and wrote the result to an external document in four seconds as one intent.

> **The question to ask.** Show me, right now, every agent in our environment, the apps each can access and the last full session one of those agents ran, call by call. An answer that needs a week and a log-aggregation project is itself the answer.

### Authentication and credentials: the agent should never hold the real key

This dimension asks what the agent holds while it runs, because that is what an attacker obtains if the agent is compromised. The usual route to compromise is prompt injection, where instructions hidden in a document, a web page or a code comment redirect what the agent does next.

**How it fails.** The default is that the agent holds live credentials sitting in the environment of the process running it. In April 2026 Aonan Guan, with Zhengyu Liu and Gavin Zhong, planted instructions in GitHub pull request titles and issue comments. The coding agents running as GitHub Actions read that text as part of their task, then read their own environment variables and posted [API keys](/glossary/api-key) back into the repository in comments and commits.[11] Any credential an agent can access can be steered out of its environment by data the agent was asked to read.

> **The question to ask.** If this agent's context is completely compromised right now, what credential does the attacker walk away with and how long is it valid? An answer naming the employee's own token, valid until somebody rotates it, describes a compromise that turns into an incident.

[_The guide_](https://www.cakewalk.security/ciso-guide)_ separates delegation from impersonation as a design choice and sets out the Comment and Control mechanism in full._

### Authorization and policy: evaluate every action as it happens

Authorization decides whether a specific action may run. A company authenticates a person once and trusts their judgment for the rest of the session. An agent has no judgment to trust, because the data it processes can steer it. The control point therefore moves from the login to each individual action.

**How it fails.** The PocketOS agent from the first chapter shows this failure. Its session already allowed deletion, so nothing stood between the agent's decision and the deleted database. A quieter failure is governing an agent with another agent, because a model evaluating a model inherits the weaknesses of the thing it judges, including the same susceptibility to injected instructions. A rule the agent can be convinced to skip is not a control.

> **The question to ask.** When this agent tries to delete a production resource, what evaluates that specific call before it runs, is the decision the same every time and can the agent's own reasoning influence the outcome?

### Human oversight: interrupt only for the decisions that change the outcome

Human oversight is a person deciding whether a specific action proceeds before it runs. It is the one control here that depends on a person's attention rather than on configuration, which is why it degrades exactly when it is needed most. A person asked to confirm every agent action soon starts approving without reading.

> **The question to ask.** How many approvals will a reviewer see in a busy day and what stops a prompt injection manufacturing one they wave through? If every action prompts a human, you have built a rubber stamp and told your team to lean on it.

[_The guide_](https://www.cakewalk.security/ciso-guide)_ explains why approval fatigue is exploitable and how binding an approval to the exact proposed action stops a stale decision executing._

### Audit: record every action in terms you can defend later

Visibility is about seeing a session while it happens. Audit is about answering for it afterward, to an incident responder at 2 a.m. or to a regulator after something went wrong. It fails on coverage, where a client logs the request while the tool calls that touched data go unrecorded. A second failure is altitude, where the log captures individual calls and never the session as a connected unit.

> **The question to ask.** Take one agent session from last week and show me every action, the person behind it and the policy decision at each step, exported into our monitoring platform. If any of it lives only in an agent client, the audit trail has a hole exactly where an incident will be.

[_The guide_](https://www.cakewalk.security/ciso-guide)_ spells out what each record has to hold for an investigation to reconstruct a session._

### Deployment and data residency: know where the governing layer runs and what it keeps

The first five dimensions govern the agent. This one covers the governing layer itself, because any control point between agents and the apps they access is a data processor that sees every prompt and payload passing through it.

> **The question to ask.** Where is an agent request evaluated, what does the layer keep after it decides and can we run it entirely inside our own infrastructure if our data classification requires it?

[_The guide_](https://www.cakewalk.security/ciso-guide)_ works through the _[_sub-processor_](/glossary/sub-processor)_ chain under _[_GDPR_](/glossary/gdpr)_ Article 28(2),_[12]_ why the regulation counts evaluation as processing rather than storage and how cloud and self-hosted deployments compare._

### Lifecycle: govern the agent from creation to decommissioning

Human identity governance has always had a lifecycle at its center, joiner, mover and leaver, because access nobody reviews or revokes is where risk accumulates quietly. Agents have one too, missing every event that governance relies on, which leaves the access an agent is granted at connection as the access it holds until somebody narrows it.

> **The question to ask.** For a given agent, who created it, what access was it granted at creation and why that much, when was that access last reviewed and if we needed to revoke everything it can do right now, how long would that take?

[_The guide_](https://www.cakewalk.security/ciso-guide)_ describes what registration and revocation look like for an actor with no joiner, mover or leaver event._

## Where Enforcement Sits Decides What It Can See

The seven dimensions describe what good governance does, while the place in the stack where it runs decides what it can observe at all. In February 2026 Herman Errico published Autonomous Action Runtime Management, an open specification for securing AI-driven actions at runtime.[13] It argues that the security boundary for an agent is the action boundary, the moment it tries to act on an external system. It proposes four implementation architectures with different trust properties.

### How the Four Enforcement Architectures Compare

_Each sits at a different layer, which is why each is blind to something the others can see._

A protocol gateway sits on the wire and evaluates every tool call that crosses it, which leaves it blind to any action an agent takes outside the protocol. SDK instrumentation reads the model's intent because it runs inside the process it polices, which also means whoever controls that process controls the instrumentation. Kernel enforcement through eBPF resists tampering, at the cost of seeing system calls rather than what they were for. Vendor integration is the most precise but reaches only as far as the vendors who have built it. The blind spots differ, which is why a serious program runs more than one at the same time. [Governing the New Frontier](https://www.cakewalk.security/blog/governing-the-new-frontier-the-missing-layer-for-agent-access) works through each architecture and what it takes to turn one of them into governance.

**Why this place in the stack.** Cakewalk builds the protocol gateway, because the other three architectures depend on something a company may not be able to change: the agent's own code for SDK instrumentation, privileged access to every host for kernel enforcement and a vendor having built the integration. A protocol gateway depends on the connection between an agent and the apps it uses, which the company already operates. That connection is also where the irreversible actions pass, because deleting a record, sending a message or moving money all cross it as a tool call.

**What it does.** Policy is evaluated on each tool call before the call executes, so the decision is made while the action is still a proposal rather than after it has run. [Credential mediation](/glossary/credential-mediation) keeps the real token on the gateway instead of in the agent, which is what leaves a compromised agent with nothing worth stealing. Every call leaves an audit record of the action, the person behind it and the decision that let it through.

The gateway also keeps the session's history, which is what lets it judge a chain rather than a single call. An agent allowed to read the customer database and allowed to send email may do both. Reading customer records and then mailing them to an outside address is [exfiltration](/glossary/data-exfiltration), which only a control that remembers the first call can recognize.

### How a Tool Call Passes Through the Gateway

_Because the credential stays on the gateway, the agent never holds it and each call is decided before it runs._

**What it does not cover.** The gateway sees what crosses the protocol and nothing else. It cannot see the agent's reasoning, its memory or an action that never becomes a tool call, such as a local file read or a shell command. Its policy binds to the tool and the action type rather than to an individual resource, which makes "delete a production resource" a rule it can express and "delete this particular database" one it cannot.

## Start Where the Irreversible Actions Are

The dimensions can read as a demand to fix everything at once. No team works that way, which is why two rules order the work. Establish visibility before enforcing anything, because an inventory of which agents exist is the fastest way to turn a vague worry into a problem somebody will fund. Then govern the actions that cannot be undone before the ones that can, because that is where one wrong call becomes unrecoverable and where credential mediation removes the standing secret that turns a compromised agent into a breach. Human oversight, audit, deployment posture and lifecycle follow, tightened as the agent footprint grows.

> **Where to start.** The same seven dimensions are scored in [Cakewalk's gap analysis](https://www.cakewalk.security/gap-analysis), which produces a prioritized view of where a company's exposure is largest. That view turns the ordering principle above into a sequence for one environment.

These seven dimensions cover one layer, the access and action boundary where an agent's permissions meet the systems it can affect. Three related risks sit outside that layer and need their own controls: a model talked into producing harmful content such as malware, false documents planted in the sources an agent looks up and a malicious package or [MCP server](/glossary/mcp-server) installed in the agent's tooling.

An agent security program is measured by whether a security team can answer, for any agent in its environment, a short list of questions with evidence rather than hope. What is this agent allowed to do? Whose access does it spend? Who approved the consequential actions? What did it do? Can its access be revoked in full, on demand?

Agents are going to act inside your company's systems either way. Governance decides whether you can account for what they did.

> **Read the full report.** This piece summarizes [A CISO Guide to AI Agent Security](https://www.cakewalk.security/ciso-guide), which works through all seven dimensions at length and carries what to build for each one.

## References

1. Gartner, "Gartner Predicts 40% of Enterprise Apps Will Feature Task-Specific AI Agents by 2026, Up from Less Than 5% in 2025," press release, August 26, 2025. <https://www.gartner.com/en/newsroom/press-releases/2025-08-26-gartner-predicts-40-percent-of-enterprise-apps-will-feature-task-specific-ai-agents-by-2026-up-from-less-than-5-percent-in-2025>
2. Gartner, "Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027," press release, June 25, 2025. <https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027>
3. Aim Labs, "Breaking down 'EchoLeak', the First Zero-Click AI Vulnerability Enabling Data Exfiltration from Microsoft 365 Copilot," Aim Security, June 11, 2025. <https://www.aim.security/lp/aim-labs-echoleak-blogpost>
4. Thomas Claburn, "Cursor-Opus agent snuffs out startup's production database," The Register, April 27, 2026. <https://www.theregister.com/2026/04/27/cursoropus_agent_snuffs_out_pocketos/>
5. M. Jones, A. Nadalin, B. Campbell, J. Bradley and C. Mortimore, "OAuth 2.0 Token Exchange," RFC 8693, IETF, January 2020. <https://www.rfc-editor.org/rfc/rfc8693>
6. AICPA, "2017 Trust Services Criteria (With Revised Points of Focus, 2022)," criterion CC6.2. <https://www.aicpa-cima.com/resources/download/2017-trust-services-criteria-with-revised-points-of-focus-2022>
7. ISO/IEC, "ISO/IEC 42001:2023, Artificial intelligence management system," ISO, December 2023. <https://www.iso.org/standard/42001>
8. NIST, "Artificial Intelligence Risk Management Framework (AI RMF 1.0)," NIST AI 100-1, January 26, 2023. <https://doi.org/10.6028/NIST.AI.100-1>
9. OWASP GenAI Security Project, "LLM01:2025 Prompt Injection," OWASP Top 10 for LLM Applications 2025. <https://genai.owasp.org/llmrisk/llm01-prompt-injection/>
10. OWASP GenAI Security Project, "OWASP Top 10 for Agentic Applications," December 9, 2025. <https://genai.owasp.org/2025/12/09/owasp-top-10-for-agentic-applications-the-benchmark-for-agentic-security-in-the-age-of-autonomous-ai/>
11. Aonan Guan, Zhengyu Liu and Gavin Zhong, "Comment and Control: Prompt Injection to Credential Theft in Claude Code, Gemini CLI, and GitHub Copilot Agent," April 15, 2026. <https://oddguan.com/blog/comment-and-control-prompt-injection-credential-theft-claude-code-gemini-cli-github-copilot/>
12. Regulation (EU) 2016/679 (General Data Protection Regulation), Article 28(2), Official Journal of the European Union, April 27, 2016. <https://eur-lex.europa.eu/eli/reg/2016/679/oj>
13. Herman Errico, "Autonomous Action Runtime Management (AARM): A System Specification for Securing AI-Driven Actions at Runtime," arXiv:2602.09433, February 10, 2026. <https://arxiv.org/abs/2602.09433>
