When AI agents hold the keys: the new perimeter of enterprise security

AI agents are shifting from interlocutors to operators. Once they hold credentials, context and autonomy, the security perimeter can no longer be declared — it has to be proven.


A scenario that’s becoming common

A recent survey by Imperial College London researchers — When Agents Handle Secrets: A Survey of Confidential Computing for Agentic AI, by Forough, Kogias and Haddadi (May 2026) — sets out this problem and proposes a way forward. Its central thesis is that the threat surface of agentic AI differs materially from an isolated model call, and that current defences, focused on software, can be bypassed by an adversary with sufficient privileges — for example, a compromised cloud operator.

This piece distils the survey’s findings and brings my own technical perspective to bear on them.

Picture the following situation: an AI agent is given access to the enterprise resource planning system, the corporate inbox and a shared folder containing financial information. The objective is legitimate — automating reconciliations, drafting responses to suppliers, preparing reports, and so on. The agent delivers. The results come through.

The most important question is rarely asked: who actually controls the environment where that agent is running?

The honest answer, in most cases, is uncomfortable. The agent runs in a third-party datacentre, on an operating system no one in the organisation has audited. The defences being relied upon — encryption, multi-factor authentication, access controls — all operate at the software layer. And that layer can be silently compromised by anyone holding privileged control over what sits beneath it.

From interlocutor to operator: the qualitative shift

A distinction worth drawing — one that often goes unnoticed precisely because it seems obvious.

In isolated calls to a model — the classic LLM use case — the cycle is closed: a question goes in, a response comes back, the session ends. The attack surface is contained.

In agentic AI, the cycle opens up. The agent plans, invokes external tools, maintains persistent memory across sessions, and delegates tasks to other agents through protocols such as MCP (Model Context Protocol) and A2A (Agent-to-Agent). It accumulates sensitive context over time, holds credentials to access systems, and operates across distributed environments that no single party controls in full. There’s no conductor.

This difference isn’t gradual — it’s qualitative. An agent that retains memory across sessions stops being a tool and becomes an actor — and the auditability of an autonomous actor imposes requirements that don’t apply to a tool: it’s not enough to examine how it was built; you also have to reconstruct what it decided to do, when, and on what information. As a consequence, the threat model changes.

Five vectors that change the game

The Imperial College London survey identifies a set of attack vectors that deserve particular attention from anyone deciding on or running enterprise AI architectures.

1. Prompt injection. Manipulation of the agent’s reasoning through untrusted inputs — a document, an email, a web page — containing concealed instructions. Once the agent can act on real systems, the impact stops being merely conversational and becomes operational.

2. Context exfiltration. Leakage of the persistent information the agent has accumulated across previous sessions. This may include document excerpts, client identifiers, internal decisions — anything held in memory to sustain ongoing work.

3. Credential theft. An agent holding access to critical systems — ERP, email, databases, APIs — becomes a high-value target. Compromising one agent can grant lateral access to multiple systems without having to compromise each one individually.

4. Inter-agent message poisoning. When agents communicate as peers via MCP or A2A, the chain of trust becomes difficult to audit. A malicious message injected somewhere along the chain can propagate and influence downstream decisions without leaving a trace in traditional logs.

5. Compromised infrastructure operator. The most troubling scenario. A cloud platform administrator, or an actor who has compromised the hypervisor or the host operating system, can access the memory of the running agent. No software defence can stop this class of attack, because it operates below the layer where the defences live.

The first four vectors are already documented in public incidents. Until recently, the fifth was treated as a theoretical hypothesis. It stopped being theoretical the moment an agent was entrusted with access to regulated data.

The proposed solution: confidential computing

Confidential computing proposes shifting the perimeter of trust from the software layer to the hardware layer. It rests on two complementary pieces.

The first is TEEs (Trusted Execution Environments) — execution environments isolated at the processor level. The code and data running inside a TEE are protected not only from other processes, but also from the operating system itself, from the hypervisor, and from whoever administers the physical machine. Even an operator with root privileges cannot read the memory protected by the TEE.

The second is remote attestation: a cryptographic mechanism that allows an external party to verify that the environment running the agent corresponds, in fact, to what was announced — same code version, same configuration, same hardware. Without this attestation, isolation loses most of its value.

The survey compares six commercially mature TEE platforms: Intel SGX, Intel TDX, AMD SEV-SNP, ARM TrustZone, ARM CCA, and NVIDIA H100 CC. The last of these cannot be sold freely to China and deserves particular note: it is the first GPU-based TEE offering with sufficient scale to run inference for large language models, opening the door to scenarios that were previously impossible — protecting not only the agent’s state, but also the model itself and the data being processed.

What’s still unresolved

It would be premature to treat confidential computing as a closed solution. The survey is explicit about the challenges that remain open, and three of them may shape architecture decisions in the coming months.

The first is compound attestation across chains of agents with multiple hops. Attesting an isolated protected environment is a solved problem; doing the same across a chain of five agents that invoke one another, some on different platforms, is still active research.

The second is TEE performance on GPU at LLM scale. The performance penalties can be prohibitive for production use cases with demanding latency requirements, and the cost/benefit relationship has to be assessed case by case.

The third is more structural: there isn’t yet a complete, consolidated framework tying these hardware primitives into a coherent security infrastructure for agentic AI in production. There’s no integrated, ready-to-adopt product yet.

What this means for decision-makers

Three points are worth reflecting on for anyone defining enterprise AI architectures.

First, principles consolidated in financial environments and critical systems — separation of duties, least privilege, auditability — need to be rethought when the actor in question is an autonomous agent with persistent credentials. Controls designed for human users cannot be transposed automatically.

Second, the compromised infrastructure operator threat model may stop being an acceptable hypothesis from the moment an agent is entrusted with personal data, financial information, matters subject to professional confidentiality, or defence-related affairs. This isn’t alarmism — it’s alignment with existing regulatory obligations.

Third, the question of which architecture to adopt has shifted. Where it used to be enough to choose the model and define the data to be processed, it’s now also necessary to answer on which hardware the agent runs, and how that execution is demonstrated to an auditor. The answer to that last question has stopped being optional for certain classes of application.


Confidential computing, taken on its own, doesn’t solve the challenge of safe agentic AI. It solves a specific and important class of problems — the ones where software, however well designed, simply isn’t enough. For the rest, what’s required remains maturity in processes, clarity in threat models, and discipline in adoption decisions.

The question left open, and perhaps worth discussing in the comments: how many organisations integrating agents into their systems have the vocabulary to discuss this class of risk with their auditors?


Glossary of acronyms

  • A2A — Agent-to-Agent. Communication protocol between peer AI agents.
  • AMD SEV-SNP — Secure Encrypted Virtualization — Secure Nested Paging. AMD’s confidential computing platform, based on memory encryption at the virtual machine level.
  • AI — Artificial Intelligence.
  • API — Application Programming Interface. Programmatic interface for integration between systems.
  • ARM CCA / TrustZone — Confidential Compute Architecture (server-oriented) and TrustZone (execution isolation predominant in mobile and embedded devices). ARM’s confidential computing platforms.
  • CC — Confidential Computing.
  • ERP — Enterprise Resource Planning. Integrated business management system.
  • GPU — Graphics Processing Unit. Today used extensively for AI workloads.
  • Intel SGX / TDX — Software Guard Extensions (secure enclaves at application level) and Trust Domain Extensions (confidential computing at the virtual machine level). Intel’s confidential computing platforms.
  • LLM — Large Language Model.
  • MCP — Model Context Protocol. Communication protocol between AI agents and external tools.
  • NVIDIA H100 CC — Confidential Computing on the H100 architecture. First confidential computing offering on GPU at the scale required for LLM workloads.
  • TEE — Trusted Execution Environment. Execution environment isolated at the processor level.

Reference: Forough, J., Kogias, M., & Haddadi, H. (2026). When Agents Handle Secrets: A Survey of Confidential Computing for Agentic AI. arXiv:2605.03213. https://arxiv.org/abs/2605.03213 — PDF: 2605.03213v2.pdf

Scroll to Top