Never let the thing check itself
The model that wrote the code, or the tool that produced the measurement, will confirm its own output when asked. Confirmation is cheap and looks like verification. A check is only a check if the checker has a different source of truth than the author.
Where it bit
Whisper word timestamps were off the measured audio edge by minus 159 to plus 224 milliseconds, repeatedly; on one job the transcript was wrong by more than a tenth of a second four times. An agent classified a real false start as a transcription artifact and would have left it in; the RMS envelope on the voice-only source showed a 126 millisecond gap and settled it. Separately, reviewer agents caught an audio displacement trap, a demuxer dropping 30-frame handles, and a stale cutlist the builder had reported as clean.
The practice
Every claim gets a second reader with zero context from the first, told to find what is missing rather than confirm what is there, with one tool and a measurement. Transcripts are hypotheses; envelopes are evidence. Tests count as a second reader only if they were not written by the same session that wrote the code.
Check it
Take any "done, verified" report from an agent and ask a fresh session, with only Read, to list what the report did not check. If it comes back empty, the second session is confirming too.
Get this check as a script you can run tonight
What to do with this
If you run a business on something AI built and the checks came back with more than you expected, that is worth a conversation.
We do a free 30-minute Health Check for service businesses that want to know exactly where their biggest leaks are. No slide deck. No pitch. We ask questions, find the gaps, and tell you what we see. If there is no obvious fix, we will tell you that too.
Blinkz finds what is broken in how a business runs, then fixes it. AI only where it earns its place.