Skip to main content
All posts

When to trust an AI number (and when to throw it out)

AIdatajudgement

There's a difference between an AI that helps you and an AI you can trust, and most of the debate about tools misses it. Trust isn't a property of the model: it's a property of the claim. Some generated claims are worth acting on. Others are worth about as much as a confident guess from someone who wasn't in the room.

Here's the test we actually use.

The four-question test

For every number or claim an AI produces, run these four questions. If any fails, treat the output as a draft.

1. Does it cite a source you can open? Not "based on public sources": a specific source. A paper, a filing, a database, a document. If the citation is vague, assume it's a hallucination until you see the source. This single habit catches more bad AI output than any other, because fabricated citations are the most documented failure mode (see why AI fails in consulting).

2. Can you reproduce it? Re-derive the number from the inputs. If the AI says a market is $2 billion, you should be able to see the arithmetic: so many customers, times so much revenue each. If you can't rebuild it, you can't defend it, and if you can't defend it, the client will find the seam.

3. Does it change your decision? This is the one most people skip. A number that wouldn't change your recommendation either way doesn't need to be perfect: it needs to be in the right ballpark. Only the decision-driving numbers need high confidence, and that's where you spend your verification effort.

4. Would you bet a client relationship on it? If the answer is yes, verify it against the primary source yourself. If it's no, label it as an estimate and move on. Most AI errors only matter when they sit in a slide that a client's board reads. Knowing which claims those are (before they're on the slide) is the whole skill.

What to trust by default

Some categories of AI output are safe enough to use almost as-is:

  • Structure and framing. An outline, a storyline, a framework: these are drafts to react to, and errors here are cheap. A wrong outline costs you an hour of rethinking; a wrong number can cost you the engagement.
  • Summaries of documents you provide. If the AI read your file, its summary is grounded in the text. Still spot-check, but the hallucination risk is low because it's working from what you gave it, not from memory.
  • Synthesis of your own data. When the model works from your spreadsheets and documents, it's mostly assembling facts you gave it. The risk here is omission and emphasis, not invention: so read it for what's missing, not for what's made up.

What to distrust by default

  • Numbers from general knowledge. Market sizes, growth rates, benchmarks pulled from the model's memory. This is the #1 hallucination zone. It's also the #1 reason AI fails in consulting. The model is predicting a plausible figure, not retrieving a real one.
  • Citations. Unless you can open the cited source, assume it's fabricated. Fake citations are the single most documented AI failure mode: from the lawyer who cited six invented court cases to the Big Four reports caught with non-existent academic references.
  • Anything stated with unusual confidence. The model's confidence tone doesn't correlate with accuracy. Suspiciously specific precision is often invented precision. A number that "should" be fuzzy and instead comes out to three decimal places deserves extra suspicion, not extra trust.

The practical rule

Use AI for the parts that are cheap to fix and hard to start: structure, first drafts, synthesis. Verify the parts that are expensive to get wrong: the numbers that drive the decision, the citations that carry the argument.

That's the whole discipline. It's not about trusting or distrusting AI. It's about knowing which claims earned verification and which ones didn't.

Why this gets harder, not easier

There's a trap worth naming: the models keep getting more fluent, and fluency reads as reliability. As the tone gets more polished and the citations get more plausible, the confidence gap between the output's surface and its underlying certainty widens. The four-question test doesn't get easier with better models: it gets more necessary.

This is also why the tooling question matters. The setups worth paying attention to are the ones where verification is built into the workflow rather than left to the reader. If a platform attaches a source to every claim as a matter of construction, the "can you open it" question is answered by design, not by your vigilance.

A field guide to the specific tells

Beyond the four questions, there are patterns that show up so reliably they're worth memorising:

  • The too-perfect citation. Author, journal, volume, page numbers all present, and the paper doesn't exist. Real citations usually carry small imperfections; fully-formed ones are often hallucinated.
  • The round number. Markets sized at "exactly $5 billion" or "2.5 million customers" without an arithmetic trail. Real estimates come with messiness and a range.
  • The hedge that isn't. "Research suggests..." or "studies indicate..." with nothing behind them. It reads like caution but it's an admission the model is reaching.
  • The outdated-as-fresh detail. A number that's confidently stated as current when the real figure moved years ago. The model isn't checking its watch; it's predicting what a current answer looks like.

None of these are proof of a hallucination: they're reasons to verify. The verification itself is the only thing that settles it, and it's cheap: open the source, re-derive the number, decide.

The mental model that makes all of this automatic

The fastest way to internalise the whole test is to stop thinking of the AI as a source and start thinking of it as a process. You don't trust a process: you audit it. Every answer is a draft produced by a system that is fluent by design and accurate only by accident of training data. Once that reframe sticks, the four questions stop feeling like a checklist and start feeling like common sense: of course you open the source, of course you rebuild the number, of course you only sweat the claims that would change the decision.

The teams that get burned by AI are rarely the ones that used it carelessly. They're the ones that treated a process as if it were a person, and a very confident person at that. The teams that get value from AI treat every output the way they'd treat a strong junior's first draft: gratefully, but with a red pen.

If you want to see a setup where every claim carries a source by construction (so the verification is built in rather than bolted on) look at the interactive demo, or try the SWOT generator and notice how it asks you to supply the context it builds from.

Get new posts and free tools

Join the list: one email when we publish. No spam.