Insights · Enterprise AI agents

Enterprise AI agent security

Chris Nielsen, PhD ·

The security problem with AI agents is not that models leak data. It is that agents hold permissions and take actions, and the instructions that drive those actions can arrive inside the content the agent reads. An attacker who cannot reach your database directly may be able to reach it through a document your agent retrieves.

That is the core of the threat model, and it is genuinely different from the threat model for a chatbot or a conventional application. This article covers the failure modes specific to agents, the controls that work, and the ones that only look like they work. For the wider deployment picture, see enterprise AI agents.

The agent-specific threat model

Indirect prompt injection

The defining agent vulnerability. An agent retrieves a document; the document contains text crafted to look like instructions; the agent acts on it. The attack does not require access to your systems — it requires the ability to get content into a corpus your agent reads. What makes this hard is that there is no reliable way to distinguish instruction from data inside a text stream. The consequence for design: injection cannot be solved at the model layer. It has to be contained at the permission and action layer, so that a successfully injected instruction cannot do anything worth doing.

Over-privileged tool access

The most common and most consequential architectural mistake. An agent is given a service account with broad read access because it is simpler than wiring per-user permissions. The result is that any user who can invoke the agent can potentially reach any data the agent can reach — and it will not show up in a conventional access review because on paper the user's permissions never changed. The control is straightforward and frequently skipped: the agent acts with the requesting user's identity and permissions.

Excessive agency

An agent with more capability than its task requires — a drafting agent that can also delete, a retrieval agent that can also send email. Each unnecessary capability is a way for any other failure to become worse. The control is ordinary least privilege, applied per agent and per workflow.

Data aggregation and inference

An agent that queries several systems can assemble a picture no single system exposes. Each individual query is authorised; the combination may disclose something the access model never intended — re-identification from combined quasi-identifiers being the obvious healthcare case. This needs purpose limitation and de-identification enforced at the retrieval layer rather than trusted to the model.

Memory and context leakage

Where an agent maintains state across sessions or users, that state is a disclosure risk. Vector databases deserve specific attention: embeddings derived from sensitive documents are sensitive, and a shared index without per-user filtering will return content the requesting user is not entitled to see. Retrieval must be permission-filtered at query time, not after.

Controls that work

Bound the action space. Register tools, version them, permission them per agent and per user. This single control converts an unbounded security problem into an analysable one.

Make consequential actions require a person. Anything irreversible, outward-facing or consequential stops for human authorisation. This is the control that makes injection survivable — a successfully injected instruction reaches a human checkpoint and stops there.

Treat all retrieved content as untrusted. Every document the agent reads is potential attack surface, including internal ones. Assume containment will sometimes fail, and design so that failure is not catastrophic.

Remove egress where the workflow allows it. An agent that cannot reach the internet cannot exfiltrate to it — for highly sensitive workflows, a zero-egress deployment eliminates a whole category of risk rather than mitigating it.

Log at the action layer. Every tool call, with parameters, identity, timestamp and result, to an immutable store — one of the few places where compliance and security spend the same money.

Where Eclypse sits: registered, versioned modules with per-user permission inheritance, and role-based access with immutable audit logs by default. See the security posture in full on the Security & trust page.

Security is usually where agent deployments stall, and usually for architectural reasons rather than incidental ones. See how AI agents work for where these controls sit in the execution loop.

FAQ

Common questions about AI agent security.

What is the biggest security risk with enterprise AI agents?

Over-privileged tool access combined with indirect prompt injection. The mitigation is per-user permission inheritance and human authorisation on consequential actions.

What is indirect prompt injection?

An attack where malicious instructions are embedded in content the agent retrieves — a document, a web page, an email — rather than typed by the user. It requires no access to your systems, only the ability to place content where your agent will read it.

Can prompt injection be prevented completely?

No. There is no reliable way to separate instruction from data within a text stream. Robust designs contain the consequences instead of trying to prevent it entirely.

Should AI agents have their own service accounts?

Generally no. An agent should act with the requesting user's identity and permissions, so it cannot reach data that user could not reach directly.

Show us your deployment constraints — we'll tell you what the threat model actually looks like.

Book a working session
ECLYPSE AI

Our proprietary AI orchestration platform for healthcare: one engine, a registry of reusable task modules and domain agents, and a governed knowledge base — deployed inside your walls and run by your team.

© 2026 Eclypse Pte. Ltd. Privacy policy
Singapore