Every few months there's another headline about AI quietly making things up. The scary part isn't the headline: it's that almost every one of them was caught late, after money had already changed hands.
Let's look at the ones that actually have paper trails, because they tell you exactly where the failure happens.
The cases that should make a consultant nervous
Deloitte Australia, A$440,000. In July 2025 Deloitte handed the Australian government a report. Three months later it was discovered to contain fabricated academic sources and a fake quote attributed to a federal court judgment. Deloitte revised the report and issued a partial refund. This wasn't a rogue intern with a chatbot: it was a Big Four firm's paid deliverable.
Deloitte Newfoundland, CA$1.6 million. The following month, a news investigation found that a Deloitte health-human-resources plan for the Government of Newfoundland and Labrador contained at least four citations to research papers that don't exist. A $1.6 million report, with fake footnotes, caught by journalists rather than by anyone who owned the work.
Mata v. Avianca. Back in 2023, a lawyer filed a brief citing six court cases that never existed, because ChatGPT invented them. The judge was not amused. A federal court in Texas went further and banned AI-generated filings that hadn't been reviewed by a human. In the years since, academic researchers have logged well over a thousand documented instances of hallucinated citations in legal decisions alone.
Air Canada. A passenger asked the airline's chatbot about bereavement fares. The bot invented a policy that didn't exist. When Air Canada tried to argue the chatbot was a "separate legal entity responsible for its own actions," the tribunal rejected it and ordered the airline to honour the made-up policy.
None of these were technical edge cases. They were ordinary uses of an LLM doing exactly what LLMs do: generating text that sounds confident.
Why this keeps happening
Here's the uncomfortable truth about how these models work. A language model isn't answering your question. It's producing the most probable next string of words given everything it's seen. It has no internal ledger of "true things" and "false things": it has probabilities. When it doesn't know, it doesn't stop and say "I don't know." It produces a fluent answer anyway, because in the training data, questions are usually followed by answers.
That's why a tool like ChatGPT can be described as "an eager-to-please intern who sometimes lies to you." The mechanism that makes it brilliant at summarising also makes it brilliant at sounding certain.
Three specific failure modes show up over and over in consulting work:
1. Invented citations. The model produces a reference that looks perfect (author, journal, year) but the paper doesn't exist. This is the Deloitte failure and the Mata failure. It's the most dangerous one because a footnote looks like a fact.
2. Confidently wrong numbers. Ask for a market size, a growth rate, or a cost figure and the model will happily produce one. It has no way to know the real figure; it's predicting what a plausible figure looks like.
3. Blended facts. The model mixes real information with invented detail so smoothly you can't separate them. The headline is true, the year is wrong, the source is someone else. This is the insidious one: you can't spot it by reading for plausibility, because it's mostly plausible.
Why consulting is uniquely exposed
Consulting is a profession that runs on the authority of the written word. A slide with a citation is treated as grounded. A model that fabricates a citation exploits that trust directly: it's not generating a hallucination, it's generating the appearance of rigour, which is worse.
The economics make it worse. Firms are under pressure to use AI to compress engagement time. When an analyst hands a deck to a manager and says "the model verified the sources," the verification step is where the whole chain lives or dies. If that step is a quick glance instead of an actual check, the client is one presentation away from discovering the error on their own: which is how a $440,000 report ends up in the news.
The checks that actually work
You can't stop a model from hallucinating. You can only stop a hallucination from surviving your process. These four checks catch the overwhelming majority:
1. Verify every citation against the actual source. If the report cites a paper, open the paper. If it cites a court case, look the case up. This catches the Deloitte failures and the Mata failure in about ten minutes per citation. There's no substitute: it has to be done against the real source, not against another AI summary of the source.
2. Require a source for every number. A market size, a growth rate, a cost benchmark: each needs a named source or an explicit assumption. "Analyst estimate" is not a source. This single rule kills most of the plausible-but-wrong numbers.
3. Run the "reverse summary" test. Have the model summarise its own answer back to you, then check the summary against the answer. Errors that survive a first pass often don't survive the second.
4. Build the answer from your documents, not from memory. An AI that has your actual engagement documents (the client's data, your market research, the source files) can ground its output in what you gave it. An AI answering from general knowledge is answering from what it thinks the world looks like, which is exactly where the fabrications live.
The uncomfortable conclusion
The failure mode isn't going away. Models get better at sounding right far faster than they get better at being right. The firms that win aren't the ones that ban AI: they're the ones that treat AI output as a draft from an overconfident junior, subject to the same verification that any junior's work already goes through.
That's the whole game: keep the speed, add the proof.
Want to see what that looks like in practice? The interactive demo shows a deliverable built with sources attached to every claim, and the free market sizing tool lets you generate a market estimate that forces you to write down the assumption next to each number.
Get new posts and free tools
Join the list: one email when we publish. No spam.