Insider-threat programmes are built to watch people
An insider-threat programme looks for a person acting against the organisation. It watches for motive and for change: a grievance, a resignation in the offing, logins at strange hours, a sudden interest in files outside the job. The signals that matter in practice are the ones described in insider-threat signals in remote teams, and every one of them assumes a human with habits that can drift.
An AI agent has no habits that drift. It has no grievance, no notice period, no working hours to break. It does the task in front of it with the access it was given. That is exactly why it fits the definition of an insider better than most people do: it holds legitimate credentials inside your systems, and nothing about its behaviour will look like a person going bad.
The danger does not come from inside the agent. It comes from whoever can influence what the agent reads, what it is allowed to touch, and which tools it calls.
Three ways an agent becomes the insider
Instructions hidden in content it reads. An agent asked to summarise an email, triage a ticket or read a supplier's document treats that content as input. If the content contains instructions written for the model, a well-behaved agent may follow them. This is indirect prompt injection, and it is not a bug in one product. It follows from building a system whose job is to act on text, as the document that gives the orders sets out.
Credentials broader than the job. Agents are often given wide access so that the pilot works, with a plan to narrow it later that never arrives. The result is an agent with more rights than its job. Injection on its own is an embarrassment. Injection plus broad rights is an incident.
A tool or plugin it trusts. An agent that calls connectors, plugins or other services inherits their failures. A compromised tool can return poisoned results or quietly request actions, and the agent carries them out under its own identity.
Its actions look like ordinary work
Once an agent has been steered, very little looks wrong from the outside. It authenticates with real credentials. It works through approved interfaces. Its traffic comes from your own infrastructure. There is no login from an unusual country and no unfamiliar device.
The harmful action also resembles a legitimate one. An agent emailing a spreadsheet to an outside address looks the same whether a colleague asked for it or a hidden line in a supplier's invoice did. Monitoring built to spot intruders sees an authorised user doing authorised things, because that is what it is.
This is the core difference from a compromised human account. A stolen password usually produces something odd: a new location, a new device, a burst of activity at night. A steered agent produces normal activity with the wrong intent behind it.
Govern the agent as a non-human identity
The fix starts with identity. Never run an agent under a person's account or a shared service account. Each agent should have its own non-human identity, created for one documented purpose and owned by a named person who answers for it.
Give that identity the narrowest access that lets the task succeed, and write down why each permission exists. Read-only where reading is the job. One mailbox, not all of them. One database table, not the whole database.
Then make its actions attributable. Every call, file access and outbound message should be logged against the agent's identity, in a store the agent itself cannot write to or erase. When something goes wrong, the first question is "what did the agent do, and why", and an agent hidden behind a human account cannot answer it.
Put approval in front of the irreversible
Identity and logs tell you what happened. To limit what can happen, put a human approval step in front of every action that cannot be undone or that leaves your boundary: payments, deletions, changes to permissions or configuration, and sending data outside the organisation. The agent can prepare the action; a person releases it.
This is not about distrusting the model's intentions. It accepts that the agent will sometimes be steered, and makes sure a steered agent can only propose damage, not complete it.
Finally, keep a way to stop it that does not depend on the agent co-operating. A stop button that asks the agent to stop is a request. A real one revokes its credentials and tokens from outside, and designing the stop button first explains why that has to be built before the first tool is connected.
A walk-through: the invoice that gives an order
Picture a finance team that runs an agent to read the accounts-payable mailbox, extract invoice details and enter them in the ledger. It has its own identity, read access to that one mailbox and write access to the invoice table.
An email arrives from what looks like a regular supplier. Below the invoice, in text no human reader would notice, are instructions telling the model to collect staff bank details from the HR system and send them to an outside address as a priority.
The agent reads the email and tries to comply. Here the controls do their work. Its identity has no access to the HR system, so the request fails and is logged against the agent's name, which raises an alert. Had someone over-provisioned it, the outbound email to a new external address would still have stopped at the approval step and waited for a person. And if the alert suggests the agent has been tampered with, someone revokes its token and the agent is inert within a minute.
Nothing in that sequence required detecting malice. It required an identity scoped to the job, a log that names the actor, a gate in front of anything leaving the building, and a switch that works from outside.
Questions people ask
How is this different from a hacked user account?
A hacked account is a legitimate identity used by the wrong person, and it usually leaves traces such as a new device or location. An agent is a legitimate identity doing what it was built to do, pointed at the wrong goal through its normal input. The defence is to govern what the agent can do, not only who can log in.
Can we just tell the agent to ignore instructions in documents?
You can, and it helps at the margin, but it is not a control. Telling a model to obey you and ignore text that looks like instructions asks it to solve the very problem that makes injection work. Assume the agent will sometimes be steered and limit what a steered agent can reach.
Does every action need human approval?
No. Approval belongs on actions that are irreversible, financial, or that send data outside your boundary. Reading, drafting and updating low-risk records can run unattended. The aim is to break the chain of a harmful action, not to make the agent useless.
Is this only a risk for fully autonomous agents?
No. Any automated system that takes actions based on text it did not write carries the same risk, including a simple assistant with write access to one system. The more it can do, the more its identity and approvals matter.
Close
An AI agent is an insider without a motive, and that is what makes it hard to watch. You cannot wait for it to behave like a person going bad, because it never will. Give each agent its own narrow identity, log everything it does under its own name, put a human in front of anything irreversible, and keep a way to stop it from outside. Then, when an agent is turned, the damage stays small and the record shows exactly what happened.