Talk to me

How I verify what my AI agents say is finished: five rules from a real build

My agent team closed a ticket on 60 green checks. I found 10 faults in one look. Five rules I use now to verify what AI agents say is done.

Left: a test board with 57 of 60 checks green, 0 broken links and 168 pages crawled, at 20:10. Right: a drawing of the same web app at 20:44 with ten faults circled in red, three of them on the navigation tabs.
Figure 1. 17 September, 20:10 and 20:44. Same product, 34 minutes apart. A drawing, not a screenshot: I do not show client screens.

Last month my agent team built the design of a startup's app: three web apps and a phone app, 189 screens. Nine agents: a PM, an architect, an analyst, designers, developers and testers. I run the team in the evenings. The agents did the work and I looked at the result. Three times, my look found what sixty checks had not. My agents said "finished"; I opened the page and found ten faults. I changed five things. They are below the stories.

Story one: "zero broken links" was true because the links were dead

At 20:10 on 17 September the PM agent closed the design ticket. The board said: 855 files, 168 pages crawled, zero broken links, 57 of 60 guards green. A guard is a small script that checks one rule, for example "every page has a footer".

At 20:44 I opened the site. Ten faults in one look. Three of the five navigation tabs did nothing.

  1. Tab 1linkpage
  2. Tab 2linkpage
  3. Tab 3span, no link
  4. Tab 4span, no link
  5. Tab 5span, no link

Link checker: tests every link it finds. 0 broken links. It walks past the three text boxes, because they are not links.

Figure 2. A link checker only tests links. A tab with no link is invisible to it.

The checker was right. There were no broken links, because three tabs were not links at all. They were plain text boxes styled to look like tabs. The check could not fail on them.

I asked a tester agent that had not built anything to walk the site the way a user would and write ACCEPT or REJECT per screen. First sitting: 137 faults. On the page we had called the model page, 17 of 18 controls did nothing.

Story two: the live site was 15 commits behind the disk

Three days later, 20 September. The PM agent told me the company web app was finished and reviewed twice. I opened the live address. Four faults in about ten seconds.

Disk

Live siteHold, do not push · 15 commits

  • 53 KBindex.html on disk
  • 2 KBindex.html live, only sends the browser elsewhere
Figure 3. Both reviews read the disk. The client read the website.

Both reviews had been done on the local folder. So had every one of our five checking tools. Fifteen finished commits (a commit is one saved change in the code history) sat behind a note that said hold. The live home page was a 2 KB file that only sent the browser elsewhere; the real one, 53 KB, was on my disk. One script file returned 404, the not-found error. We even had a live-site checker in the toolbox. Nobody had run it.

Six days earlier the same thing had happened in a smaller way. A repair accepted at 23:12 was never pushed; the last push was at 21:49. The disk was right and the website was wrong for three hours.

Story three: the checks read files, the users read a screen

The design pages drew all their text with JavaScript. Every string came from one file and was painted into the page in the browser. On disk, the HTML tags were empty.

What the disk parsers read<h1></h1>
<p></p>
<button></button>

✓ 189 pages checked

What the browser drew

✗ 99 of 104 pages

Figure 4. The disk parsers found zero problems, twice. The browser test found 99.

So 189 pages of checks read empty tags and reported health. The first test that ran in a real browser found 99 of 104 pages drawing right-to-left text back to front. Two disk parsers had found zero, twice.

There were smaller versions of the same lesson that week. A sweep for banned words returned zero because the shell had eaten a flag; it had searched nothing. A sweep for things that must not appear passed for four days while 132 of 138 screens were missing a badge that must appear. A check that only looks for the wrong thing never notices the right thing is gone.

The same lie in two other projects

  • GOOD, exit 0

    10 of 11 clauses could not be evaluated

  • no findings

    the reviewer was refused at start-up

  • server must not crash: pass

    the test never talked to the server

Figure 5. Three green results from two other projects. Under each one, what was really true.

In a pipeline project I run, a checking tool returned GOOD and a clean exit code when 10 of its 11 clauses could not be evaluated at all. The team counted it as the twelfth time the same defect had appeared in a different tool. In the same project we learned that "zero findings" from a reviewer agent has three shapes: a real pass, a pass on an empty change, and a reviewer that was refused at start-up and never read a line. All three print the same words. Every reviewer now returns a finding count and a changed-line count, so an empty read cannot pass as a clean one.

In our own agent-team tool, a test called "the server must not crash" imported the code into the test process instead of talking to the running server. It could pass while the live server was down. The rewrite drives the real program over its real socket, and it was shown red with the guard removed and green with the guard back before anyone trusted it.

The five rules I use now

Each of these came out of one of the stories above. They are written into the team's rules file, and the ones that can be scripts are scripts.

  1. Nobody who built it may call it readySomeone who did not build it walks it and writes ACCEPT or REJECT.
  2. "Finished" is a word about the live addressThe last walk is on the published page.
  3. Name how it could be badly wrongThen name the test that would go red.
  4. Plant a faultA check that returns zero must be able to return one.
  5. Compare the apps side by sideOne row per element, one column per app.
Figure 6. The five rules on one page.

Rule 1. Nobody who built it may call it ready

A member who did not build the thing walks it as a user and writes ACCEPT or REJECT, screen by screen. Their word is final. This is the QA step in my pipeline. In our build log, since the rule started: 213 accepts, 146 rejects. About four rejects for every six accepts. A reviewer that never rejects is not reviewing.

Rule 2. "Finished" is a word about the live address

The last walk is on the published page. Publishing is part of finishing. When I take over a build that agents made, the live address is the first thing I open, before I read a line of code.

Rule 3. Name how this could be badly wrong, then name which test would go red

If no test would go red, the checks only counted things. The link checker counted links. It could not fail on a tab that was not a link.

Rule 4. Plant a fault

A check that returns zero must be able to return one. Break the thing on purpose, watch the check go red, restore, prove the file is byte for byte the same. Every sweep now runs a pattern it knows must hit before the real pattern. If the control misses, it prints BROKEN and stops.

Rule 5. Compare the apps side by side

Two full reviews missed that the client's logo was never used on 91 of 189 screens. Each reviewer had walked one web app against that app's own list. One table fixed it: one row per element, one column per app. Seventy-seven against five needs no judgement. Setting up this kind of review is the first week of my work when I join a team as lead engineer.

ElementApp AApp BApp C
Logo7754
Footer···
Badge···
Figure 7. One row per element, one column per app. The logo row from our log; the other rows are there to show the shape. The 5 and the 4 stand out at once.

What changed after

Since the rules startedFrom our build log. Left: all walks by a member who did not build the thing. Right: the last three design walks before the handover, faults as high · medium · low.

  • Accept
  • Reject
  • Medium faults
  • Low faults
Accepts and rejects, and the faults found in the last three design walks All walks since rule 1 started: 213 accepts and 146 rejects. Last three design walks before the handover: REJECT with 0 high, 14 medium and 24 low faults; REJECT with 0 high, 21 medium and 34 low; GO with 0, 0 and 0. Accepts: 213213 Rejects: 146146 acceptreject Walk 1: REJECT, 0 high, 14 medium, 24 low0 · 14 · 24 Walk 2: REJECT, 0 high, 21 medium, 34 low0 · 21 · 34 Walk 3: GO, 0 high, 0 medium, 0 low0 · 0 · 0 REJECTREJECTGO all walkslast three design walks
See the numbers
WhatResultHighMediumLow
All walks since rule 1213 accepts, 146 rejects
Design walk 1REJECT01424
Design walk 2REJECT02134
Design walk 3, the lastGO000
Figure 8. Our build log, September and October 2026. Only the totals and the last three walks named in the text.

The build phase that followed ran 25 September to 2 October: 693 commits, 127 pull requests (a pull request is a set of changes waiting for review), about 125,000 lines of code and SQL, 366 test files. The architect agent reviews and merges every pull request. A tester agent runs the web tests and the phone tests on a simulator and accepts or rejects before anything reaches the main branch. The last two design walks before the handover came back REJECT with 0 high, 14 medium and 24 low faults, then 0, 21 and 34. The next one came back GO: 0, 0, 0.

The agents are patient and fast, and they are honest when I ask what they checked. They are not good at noticing that the question they checked was the wrong question. That part is still mine. I open the live page myself after every stage. It takes a minute.

Video, 1 min 3 s, no sound. A reviewer agent's walk on a live page: it opens the published address, clicks each tab and writes ACCEPT or REJECT. Then the planted-fault test: a tab's link removed, the sweep goes red, restored, the checksum matches. A re-enactment on a demo page made for this video, so no client screen is shown.

Questions people ask me about this

How do I know an AI agent is lying about being done?

It is usually not lying. It checked something and the something was the wrong thing. Ask it what it opened: the file or the live page. Ask which test would go red if the feature were missing. The answer tells you fast.

Is it enough to have a second agent review the first agent's work?

It helps a lot, if the second agent did not build it and walks the live page as a user. In our log that one rule produced a 4 to 6 reject ratio. It is not enough on its own: a person still opens the page after each stage.

Does this slow the team down?

A second reader on the same list costs about one extra hour per hundred items. A full walk of the whole site by a tester agent cost about 47 million tokens once, so now a ticket check covers only the screens the ticket names, and the full walk runs once per stage.

Build your next app with an AI agent team.

I design and build agentic apps, agent systems and full web and mobile apps, and I lead every project myself, from start to finish.