Insights · Agents

What should an AI agent never be allowed to do alone?

Anything you can’t undo, anything that leaves the building, and anything that changes who can do what. Here’s where we draw the line, and why telling the agent “don’t” isn’t enough.

When a chatbot gets something wrong, you get a bad answer. You read it, frown, and move on. When an agent gets something wrong, it does the wrong thing. It sends the email, issues the refund, deletes the folder. By the time anyone reads about it, it has already happened.

That is the whole shift. Once software can act, the useful question is no longer “how good are its answers?” It is “what happens on the day it’s wrong?”

That day will come. Current models are capable, but they still misread instructions, fill gaps with confident guesses, and now and then do something nobody asked for. They can also be steered by what they read. If an agent works through your inbox, anyone who can email you can put words in front of it, and some of those words will be written to look like instructions. This is called prompt injection, and nobody has a complete fix for it yet.

None of this makes agents a bad idea. We build them, and they are genuinely useful. It just means deciding in advance which actions an agent takes on its own, and which ones wait for a person.

Three questions

Before we give an agent a new ability, we ask:

  1. Can it be undone? A draft can be deleted. A sent email can’t be unsent. A payment can sometimes be recovered, slowly, after a lot of phone calls.
  2. Does it leave the building? Anything that reaches a customer, supplier, regulator or the public goes out in your name. Internal mistakes are embarrassing. External ones cost money.
  3. Could you explain it to the person it affected? If this went wrong, would you be comfortable telling them that software made the call on its own?

If the action can’t be undone, leaves the building, or would be hard to explain, a person should be involved.

The list

These are the things we would not let an agent do alone, whatever the organisation and however good the model.

Move money

Paying invoices, issuing refunds and, above all, changing bank details. Invoice fraud has worked for years with one simple move: a convincing email saying a supplier’s account has changed. An agent that reads that email and updates the record is doing exactly what the fraudster hoped a busy person would do. Let the agent prepare the payment run. Let a person release it.

Delete what can’t be recovered

Customer records, live data, backups, shared drives. Agents tend to tidy up, and they will clear out things that look old or duplicated if they think that is the job. If something needs deleting, make it recoverable first, by moving it to a bin that empties later. Keep permanent deletion for people.

Speak for you in new situations

Drafting is fine. Routine, templated messages are usually fine too: an order confirmation doesn’t need a human. The line is novelty and consequence. A first reply to an angry customer, anything to a regulator or a journalist, anything that commits you on price, timing or liability: a person reads it before it goes.

Send data outside your organisation

Uploading files to a new service, forwarding attachments, pasting records into another tool, sharing a folder link. This is how data leaks without anyone meaning it to. It is also what a prompt injection usually tries to trigger: “summarise this and send it to…”

Change permissions, including its own

Granting access, creating accounts or keys, editing its own instructions or settings, switching off logs. An agent that can widen its own powers doesn’t have limits. It has suggestions.

Make decisions about people

Hiring, dismissal, disciplinary action, credit, whether someone gets a service. An agent can gather the facts and set them out clearly. A person makes the decision and owns it. UK data protection law also has specific rules on decisions made solely by automated means that significantly affect someone, so take advice before automating anything close to this.

“Not alone” has to mean something

A checkpoint only works if the person at it can tell what they are approving. Compare these two:

Unhelpful

The agent wants to perform an action.

Useful

Pay £4,200.00 to Northgate Supplies Ltd

To account
ending 7719 new: changed 2 days ago
Previously
ending 3402
Requested by
email from accounts@northgate-supplies.co, 09:14 today
An illustration. The names and numbers are made up.

The first one asks for trust. The second gives the person what they need to spot a problem: how much, to whom, what has changed, and where the request came from.

Two more things matter.

Fewer, better checkpoints. If an agent asks for approval fifty times a day, people stop reading and start clicking. Put checkpoints where the consequences are, and let everything else run.

Enforce limits outside the model. Writing “never make payments over £500” in the agent’s instructions is a request, not a control. The agent can misread it, and an injected instruction can argue with it. Real limits live in the systems around the agent: login details that can’t reach the payments system, a spending cap set in the bank’s own platform, a fixed list of addresses it is allowed to email. If the agent can’t do something, it doesn’t matter whether it has been talked into it.

What agents can do alone

Most of the useful work, as it happens. Reading, searching, sorting, summarising, drafting, checking, preparing. Anything internal and reversible. An agent that puts together the payment run, drafts the replies, flags the odd invoice and sets out the facts for a decision saves real time, and none of it needs anyone to hold their breath.

The point isn’t to keep agents on a short lead. It’s to be clear about where the lead is, so you can give them plenty of room everywhere else.

Try this

List every tool and system your agent can use. For each action, write down two things: can it be undone, and does it leave the organisation? Anything that can’t be undone, or does leave, gets either a person in the loop or a hard limit enforced outside the agent. Everything else can run.

It doesn’t take long, and it is the quickest way we know to find the permission nobody meant to give.

Want a second opinion?

We’ll go through your list with you.

Tell us what your agents can reach. We will help you work out which actions need a person, which need a hard limit, and which can safely run on their own.