Skip to content
Go to homepageDrova logo

AI data leakage: what your people paste into AI

The fastest data breach in your business is a paste. Where it goes, why it matters, and the rules that keep the tools without the leak.

Search for "ChatGPT data leak" and most of what you find is not the AI company being hacked. It is businesses discovering that their own information walked out through an ordinary paste, by a capable person with a deadline, using the best tool they could reach.

Simple feature visual

TL;DR

  • AI data leakage is business information leaving your safeguards through an AI tool: pasted into a public chatbot, uploaded for a summary, surfaced by a connected assistant, or exposed through a compromised AI account.
  • Almost none of it is malicious. It is capable people with deadlines using the best tool they can reach, which is why it does not look like a security event to anyone.
  • The NCSC is plain about public AI tools: queries are visible to the provider, stored, and can be used to develop the service. With unapproved consumer tools, none of that sits under your settings.
  • The fix is not a ban. Bans push usage onto personal devices where you see nothing. The working pattern is approved tools on business terms, clear rules on what may go in, and a sanctioned alternative that is genuinely good.
  • On the register, AI data leakage is a driver on the confidentiality entries you already carry: customer data breach, IP loss, contractual and privacy obligations.

What is AI data leakage?

AI data leakage is confidential business information moving outside your safeguards through an AI tool. Unlike most data risks it rarely involves an attacker: the mover is your own staff, the motive is productivity, and the moment it happens looks exactly like someone doing their job well. A customer list pasted in for a summary. A contract uploaded for a plain-English explanation. A board paper dropped into a chatbot to draft talking points.

That is what makes it different from a classic breach. There is no intrusion to detect, no malware to scan for, and the person involved usually believes they did something clever. The exposure is created by where the data went and what the tool's terms allow to happen to it next.

Part of our series on the real risks of AI for a business: familiar risks, with the likelihood turned up.

The four ways out

The paste

Customer records, financials and contracts dropped into a public chatbot for a summary or a rewrite. The commonest door, and invisible to security tooling because it looks like typing.

The upload

Whole documents, spreadsheets, codebases and meeting recordings handed over for analysis or transcription. One file can carry thousands of records through the door at once.

The connection

AI assistants plugged into mailboxes, drives and CRMs can surface data to people who should never see it, and can be manipulated into sending it out: that attack has its own page in this series.

The account

AI accounts hold history: everything anyone ever pasted sits behind one login. A weak or reused password on an unapproved tool exposes months of accumulated business information in one go.

The provider's side of the paste

Where pasted data goes

What happens after the paste depends on the tool and its terms. The UK's National Cyber Security Centre puts it plainly in its guidance on public AI tools: your query is visible to the provider, stored, and can be used to develop the service, and its advice is not to include sensitive information in queries to public tools at all.

Business tiers exist precisely to change those terms: retention you configure, training exclusions, admin oversight of accounts. Which is the real dividing line in this risk. It is not AI versus no AI; it is tools whose terms your business has read and set, versus tools where an employee clicked agree on their personal account and nobody knows what was agreed to.

So the first diagnostic question is not "has anything leaked?", which is usually unanswerable. It is "which tools are in use, on whose terms?". If that list does not exist, that is the finding.

Rules for what goes in

The rules that work

Of the four families on the safeguards page, this risk belongs to the first: rules for what goes in. In practice: a data rule people can hold in their head, for example, nothing that identifies a customer, nothing under NDA, nothing you would not put in an email to a stranger. An approved-tools list on business terms, so the paste happens somewhere with retention and training settings you chose. A sanctioned alternative that is genuinely good, because a blocked tool with no replacement is how usage moves to personal phones. No punishment for asking, so the question "can I put this in?" gets asked out loud.

A ban is the tempting version and the weakest one: it converts visible usage into invisible usage while the deadlines stay. The same logic sits behind the AI policy template, which carries the data rules, the tool list and the request route as standard sections.

On the risk register, run this as a driver on the confidentiality entries you already hold: customer data breach, loss of IP, contractual and privacy obligations. The joint Engaging with AI guidance from the Australian, UK and US agencies is the reference for the broader discipline of using AI safely.

AI data leakage FAQs

What is AI data leakage?

Confidential business information leaving your safeguards through an AI tool: pasted into a public chatbot, uploaded for analysis, surfaced by a connected assistant, or exposed through a compromised AI account. It rarely involves an attacker; the mover is usually your own staff being productive.

Is ChatGPT a data leak risk?

The commonest risk is not the provider being breached. It is what your own people paste in, under terms nobody at your business has read or set. The NCSC's guidance is that queries to public AI tools are visible to the provider, stored, and can be used to develop the service. Business tiers with configured settings change that picture; unapproved consumer use is where the exposure concentrates.

Does AI train on what you type?

It depends on the product and its settings. Consumer tools can use conversations to improve the service; business tiers typically exclude training and give you retention settings. The register-level point is sharper: if the tool is unapproved, nobody at your business knows which of these applies, and that is the exposure.

Should we just ban AI tools?

Bans convert visible usage into invisible usage: the deadlines stay, so people switch to personal devices where you see nothing. The pattern that works is a memorable data rule, an approved tool on business terms, and a sanctioned alternative good enough that nobody misses the banned one.

How should our risk register handle AI data leakage?

As a driver on the confidentiality entries you already carry: customer data breach, IP loss, contractual and privacy obligations. The safeguard family is rules for what goes in, plus an approved-tools list whose settings someone has actually verified. Score each entry by asking where its data could meet an unapproved tool.

The AI Disruption Index scores the risks AI is driving against your objectives, with a safeguard drafted for each. Free, in about ten minutes.

Leaked data is just one risk. See everything AI is driving.

AI risk series

Explore related topics

The real risks of AI for a business

The series hub: what counts as an AI risk, the four families, and where to start.

Guardrails, safeguards, controls: what AI actually needs

Three words untangled, and the four families of AI-era safeguards.

The AI policy your business actually needs

The full template, free on the page, and how to make it yours.

AI hallucinations at work: examples and what they cost

Three documented cases with price tags, and the verification safeguards that catch fabrication before it ships.

Prompt injection: the attack your register hasn't heard of

Instructions hidden in ordinary content, and the safeguards that limit the damage.

AI phishing: the email with perfect grammar

Why the spot-the-typo era is over, and the safeguards that work without spotting the fake.

AI scams targeting businesses

Old cons, industrialised: the four branches, the tells that remain, and what to do if you're hit.

Fraud no longer needs a forger

Deepfakes, voice clones, invoice fraud, and the safeguards that still hold.

AI cyber attacks: when the attack is automated

Familiar attacks at a new tempo: the real uplift, the hype, and the fundamentals that still hold.

How to run an AI risk assessment

Three ways to do it, compared, and the five steps.

The AI risk register

What every entry carries, with worked examples.

AI in risk management: what it can genuinely do

The four jobs AI does well, and the three things it must never own.

What is AI disruption?

A plain definition: the change is in your risks and plans, not just your tools.

AI governance, in plain English

The four working parts, who owns what, and when a framework earns its keep.

See your own AI risk picture

The risks AI is driving against your objectives, scored for your business.