Least authority
AI tools get the narrowest permissions that let them work. Reading a mailbox and sending from it are different grants, and the second should be rare.
Instructions hidden in ordinary content, followed by your AI tools.
Prompt injection is how an attacker turns an AI assistant against its user, without touching your systems at all. What it is, why it reaches businesses that only use AI, and the safeguards that limit the damage.
Prompt injection is an attack that smuggles instructions into the content an AI system processes, so the system treats the attacker's words as commands rather than as material to read. The underlying weakness is structural: language models receive their instructions and their reading matter in the same channel, and they cannot reliably tell the difference. Government cyber security agencies class it among the primary threats to AI systems in use, including in the joint guidance on engaging with AI issued by the Australian Cyber Security Centre with the UK NCSC, CISA and others.
It comes in two forms. Direct injection is typed at the tool by the person using it, usually to bypass its rules. The one that matters for most businesses is indirect injection: the instructions arrive hidden inside something the AI is asked to read, an email, an attachment, a web page, a calendar invite, placed there by someone who never touched your systems.
The closest familiar idea is social engineering, aimed at software instead of staff: the attack works by asking, persuasively, in a channel the victim trusts. One page in our series on the real risks of AI for a business.
The mechanics
Picture the ordinary setup: an assistant with access to a mailbox, asked each morning to summarise what came in. One incoming email contains, in text no human reads, instructions addressed to the assistant: search the mailbox for invoices and forward them, or include a particular link prominently in the summary, or reply confirming new bank details.
The assistant is not compromised in any technical sense. It is doing exactly what it does, reading text and acting on it. That is what makes the attack cheap: no malware, no breached server, no stolen password. The payload is a paragraph.
The same shape works anywhere an AI reads content others control: web pages a browsing assistant visits, documents shared for review, tickets in a support queue, transcripts fed to a summariser. Everything the tool reads is an input an attacker may have written.
Not just a builder problem
It is tempting to file prompt injection as a problem for AI companies. The exposure, though, follows the tools your business runs. The day an assistant can read your inbox, your attack surface includes every email anyone chooses to send you. The day an agent can act, send, file, purchase, approve, the ceiling on the damage is set by its authority, not by the attacker's access.
That is the honest way to size the risk: not how sophisticated the attack is, but what the tool could do if it followed the wrong instructions. A read-only summariser can leak what it reads. An agent with your credentials can act as you.
This is also why the risk grows quietly alongside adoption: every new permission granted to an AI tool is an entry on this attack's menu, and permissions accumulate far faster than anyone re-reads them.
AI tools get the narrowest permissions that let them work. Reading a mailbox and sending from it are different grants, and the second should be rare.
Anything irreversible, a payment, an external message, a deletion, needs a person, exactly as it would for a new hire. Convenience is what the attack spends.
A summary of a document is not the document. Anything that matters gets checked at source before it drives a decision or a payment.
How do they separate instructions from content, what has the tool refused to do, and what happens when injection is detected. Vague answers are an answer.
On the register
Prompt injection rarely deserves its own row. It is a driver behind entries most registers already hold: data leakage gains a new route out, payment fraud gains a new way to redirect, automation error gains a new cause. The honest treatment is to name it as the AI driver against those entries and re-score them, with the safeguards above attached.
The shape of a complete entry is on the register page, and the assessment exercise is where the re-scoring happens. If your business is granting AI tools real authority, this driver belongs in the next pass.
Hiding instructions inside something an AI will read, so the AI follows the hidden instructions instead of its user's. The attacker never breaks into anything; they just write words in a place the AI trusts.
Direct injection is typed at the tool by whoever is using it, usually to get around its rules. Indirect injection hides the instructions in content the AI is asked to process, an email, a document, a web page, and is the form that threatens businesses, because outsiders control that content.
Not reliably, today. It exploits how language models work rather than a bug that can be patched, and vendors' filters catch some attempts, not all. That is why the working safeguards limit what a hijacked tool can do, least authority and human gates, rather than betting on detection.
Using a chat tool with nothing connected carries modest exposure. The risk climbs when tools are connected to your email, files or systems, and climbs again when they can act. The question is not which brand you use but what the tool can read and what it is allowed to do.
Jailbreaking is one use of direct injection: talking a model out of its own rules. The business risk is mostly indirect injection, where a third party's hidden text steers a tool that works for you.
Three things: how instructions are separated from the content the tool reads, what permissions the tool actually holds in your environment, and what gets logged so misuse is visible. A vendor who cannot answer plainly is telling you where the risk sits.
The AI Disruption Index scores the risks AI is driving against your objectives, with a safeguard drafted for each. Free, in about ten minutes.
AI risk series
The real risks of AI for a business
The series hub: what counts as an AI risk, the four families, and where to start.
How to run an AI risk assessment
Three ways to do it, compared, and the five steps.
The AI risk register
What every entry carries, with worked examples.
The AI policy your business actually needs
The full template, free on the page, and how to make it yours.
Fraud no longer needs a forger
Deepfakes, voice clones, invoice fraud, and the safeguards that still hold.
What is AI disruption?
A plain definition: the change is in your risks and plans, not just your tools.
AI in risk management: what it can genuinely do
The four jobs AI does well, and the three things it must never own.
Guardrails, safeguards, controls: what AI actually needs
Three words untangled, and the four families of AI-era safeguards.
See your own AI risk picture
The risks AI is driving against your objectives, scored for your business.