Skip to content
Go to homepageDrova logo

AI hallucinations at work: examples and what they cost

Three documented cases with price tags, why it happens, and the safeguards that catch it.

An AI hallucination is confident fabrication: output that reads as fact with no marker that it is invented. The cost is never the error itself. It is the error surviving unchecked into something that carries your organisation's name.

Simple feature visual

TL;DR

  • An AI hallucination is confident fabrication: the system produces something that reads as fact, delivered in the same tone as its correct answers. It is a property of how these tools work, not an occasional glitch.
  • The documented business costs are real: an airline held liable for a refund policy its chatbot invented, lawyers fined for filing six court cases that did not exist, and a Big Four firm repaying part of a government contract over fabricated references.
  • The pattern across all three: the cost landed when the fabrication left the building unchecked, not when it was generated.
  • The safeguard family is verification of what comes out: a named human owns anything AI-drafted that ships, citations get opened against the source, and AI never gets the last word on a fact.
  • On the register, hallucination is a driver on existing entries, wrong advice to customers, defective professional work, misleading figures, not one new IT risk.

What is an AI hallucination?

Generative AI produces the most plausible next words, not verified facts. Most of the time the plausible and the true are the same thing, which is what makes the tools useful. When they diverge, the output is fluent fabrication: a case citation, a company policy, a research paper, a number, delivered with exactly the confidence of a correct answer. Nothing in the text warns you.

Rates vary by tool and task, and grounding a system in approved content reduces them, but the behaviour is inherent to how the technology works. The practical stance for a business is to treat hallucination as an operating property to be caught, not a bug that will shortly be fixed.

Part of our series on the real risks of AI for a business: familiar risks, with the likelihood turned up.

Three examples with price tags

The invented policy

Air Canada's website chatbot told a customer booking travel for his grandmother's funeral that he could claim a bereavement discount after flying. The airline's actual policy said otherwise. In Moffatt v Air Canada, a Canadian tribunal ordered the airline to pay C$812.02, rejecting its argument that the chatbot was a separate entity answering for itself.

The six cases that never existed

A brief filed by two US lawyers cited six court decisions generated by ChatGPT, complete with quotes. Asked to verify them, ChatGPT said they were real. In Mata v Avianca, the court fined the lawyers and their firm US$5,000 for the filing.

The references to nowhere

A Deloitte report for the Australian government cited research that did not exist and included a fabricated quote from a federal court judgment. The firm repaid about A$97,000 of a contract worth about A$440,000.

The pattern

In every case the fabrication was fluent, specific and confident, and in every case the cost landed when it shipped unchecked. The failure was not using AI. It was publishing AI output nobody verified.

The surfaces that matter

Where it bites

Three surfaces carry most of the exposure. Customer-facing systems. A chatbot answering questions about policy, price or entitlement is making commitments: the Air Canada tribunal was explicit that everything on your website is yours, chatbot or not. Professional outputs. Reports, advice, citations and figures where your name is the warranty; this is where the lawyers and the consultants were caught. Internal decisions. The least visible surface: a fabricated number in a briefing shapes a real choice, and nobody outside ever learns why.

The asymmetry is what makes this risk worth a register entry: the fabrication costs the tool nothing to produce, while the unchecked version can cost a tribunal finding, a court sanction or a repaid fee. Verification effort is small next to that. The rules for who may use AI on which surface belong in your AI policy.

Verification of what comes out

The safeguards that catch it

Of the four families on the safeguards page, this risk belongs to verification of outputs. In practice: a named human owns anything AI-drafted that leaves the building, and owning means reading. Citations and quotes get opened, not skimmed. The test from the court cases: does the source exist, and does it say what the output claims? Numbers get traced to a system of record. Customer-facing tools answer from approved content or hand off to a human, and their answers are treated as commitments, because legally they are.

One safeguard that does not work: asking the AI to check itself. The lawyers tried exactly that, and the tool that fabricated the cases confirmed them. Verification happens against the real source or it has not happened.

On the risk register, run hallucination as a driver on the entries it feeds, customer misinformation, professional negligence, misleading reporting, and score each surface where AI output can ship unchecked. What AI should and should not own outright is mapped on the AI in risk management page.

AI hallucination FAQs

What is an AI hallucination?

Confident fabrication: generative AI produces plausible text rather than verified fact, and when those diverge you get invented citations, policies, numbers or quotes delivered in the same tone as correct answers. It is inherent to how the tools work, not an occasional bug.

What are real examples of AI hallucinations at work?

Air Canada's chatbot invented a bereavement refund policy and a tribunal held the airline liable for it. Two US lawyers were fined after filing six ChatGPT-fabricated court cases. Deloitte repaid part of an Australian government contract after its report cited research that did not exist. All three are documented and linked on this page.

Can you stop AI from hallucinating?

You can reduce it, by grounding tools in approved content, choosing narrow tasks and using better models, but not eliminate it. The reliable safeguard is verification: a human owns what ships, sources get opened, numbers get traced. Build for catching it, not for hoping it stops.

Does asking the AI to check its own work help?

No. In the court-citations case the lawyer asked ChatGPT whether the cases were real and it said yes. A system that fabricates fluently will confirm fluently. Verification has to happen against the real source.

How should our risk register treat hallucinations?

As a driver on existing entries rather than one new AI risk: wrong advice to customers, defective professional work, misleading internal figures. Score each surface where AI output can leave unchecked, and attach the verification safeguard to those entries.

The AI Disruption Index scores the risks AI is driving against your objectives, with a safeguard drafted for each. Free, in about ten minutes.

Made-up facts are just one risk. See everything AI is driving.

AI risk series

Explore related topics