The sorter

Can you trust what AI gave you?

One question decides how much you need to check any AI output: did you witness what it's working from, or not?

What this is for

The billion dollar question with AI is how you know whether to trust what it just handed you. Most people answer it by how important the task feels. That's the wrong measure. The real one is simpler: if it got this wrong, would you catch it?

And whether you'd catch it comes down to one thing, which is what this page is about.

Every AI job is one of two types

You witnessed it

You were there. The AI is working from something you already hold in your head, so you can check its answer against what you know. A wrong answer stands out. You are the test.

AI writes up a call you were on.

You didn't

The AI is your only source. A wrong answer looks exactly like a right one, because you have nothing in your head to check it against. You can't be the test.

AI summarises a report you never read.

The trap

Be honest about which type you reach for AI the most. It's the second one, isn't it, the stuff you didn't want to read and the things you don't know that well. That's the whole reason you handed it over.

So we end up trusting AI the hardest exactly where we can check it the least.

How much to test it

It really comes down to how much testing an AI job needs, and witnessing is the thing that decides it. If you witnessed the source, you are the test, and a careful read is enough, so you can happily let AI do more of that kind of work on its own (that's where it earns a higher rung on the autonomy ladder).

If you didn't witness it, trust isn't a plan. You have to build the checking in, because you can't be the one to catch it. Three simple ways to do that:

Make it quote its sourceHave it paste the exact lines it's drawing from, so you can spot check the claim against the words, not just trust the summary of them.
Sanity check the one numberFind the single figure the decision turns on and verify that one, at the source. You don't have to read everything to catch the thing that matters.
Ask someone who was thereIf a person witnessed what you didn't, thirty seconds of their eyes is worth more than an hour of yours. Borrow someone else's memory as the test.

Sort your own AI tasks

Paste this into Claude (or any AI you use). It interviews you about the tasks you actually hand to AI, sorts each one into witnessed or not, and hands back the single check to put in place for the ones you didn't witness. You just answer the questions.

Paste this into Claude
I want to work out how much I can trust the AI outputs I rely on, using one simple test: for each task, did I witness the source it's working from, or not. Interview me one question at a time about the tasks I actually use AI for (summarising, drafting, analysing, answering questions I don't know the answer to). Keep it easy to answer. For each one, help me decide: did I witness what it's working from (I was on the call, I know the topic, I ran the project), or did I not (a document I never read, a subject I don't really know)? If I witnessed it, tell me I'm the check and a careful read is enough. If I didn't, I can't be the check on my own, so suggest one simple way to test that output that doesn't rely on me already knowing the answer, for example making it quote the source so I can spot check, sanity checking the one number that matters, or asking someone who was actually there. When we're done, give me a simple table: each task, witnessed or not, and the one check I should put in place for the ones I didn't witness. Start with your first question.

The honest bit

  • -Witnessing lowers the risk, it doesn't delete it. Even when you were there, a summary can quietly steer you, you read its highlights and stop remembering the rest. So skim the source, don't just nod at the summary.
  • -This isn't about AI being unreliable. It usually gets it right. The point is that when you can't check it, right and wrong look identical, so you need a test that isn't just your own confidence.
  • -Making it quote its source is a spot check, not a guarantee. It can still misjudge what mattered. But it turns a summary you can't verify into words you can.

The rest of the series

This is one of a series on how to actually think about AI, not just use it. The companion is the autonomy ladder (how much to let AI do on its own), and the groundwork underneath both is getting your business AI ready.