AI
Two independent checks agreed on a number more than four times too high
A redundancy gate passed seven counts, three wrong; the worst because both paths shared one definition. Asked for an independent check, models reproduced it.
Most teams have a version of the reassurance: we checked it two ways and the numbers match. A short paper posted to arXiv on 29 September 2026 takes one such check apart, in a pipeline its own author built, and shows what the match was worth. It confirmed that two pieces of code did the same thing. It said nothing about whether that thing was right.
The study
The paper is a single-author case study by Fabio Rovai of The Tesseract Academy in London, written for a workshop titled AI for Science: Verification in the Age of AI Scientists. The pipeline compared two public catalogues of objects in Earth orbit, CelesTrak's satellite catalogue and the General Catalog of Artificial Space Objects (GCAT), and published counts of the places where the two registers disagree.
Its verification was intentionally heavy: three checks ran on every build. One was a dual-computation gate. Every published count was computed once in Python, straight from the source files, and once with SPARQL queries over a graph of 2,330,660 triples, and the build failed if any pair differed.
On the run whose numbers were published, the gate printed "ALL CROSS-CHECKS AGREE" on seven counts.
What was wrong
Three of the seven were wrong. The largest, the number of objects on which the two catalogues disagree about whether an object is still in orbit, was published as 932. After two rounds of correction, the reference figure is 220.
The cause was one line. An object counted as gone if its GCAT status code was in a set named GONE, and every other code was read as still in orbit. But GCAT's codes for docking and attachment mark the end of a phase in an object's history, not its fate. Of the 932, 769 were objects CelesTrak records as decayed and GCAT did not record as gone. Of those, 640 ended in a transition code like that, 500 of them docking or attachment.
The other two errors were different in kind. One count missed a whole file: the source lists four main catalogue files and the pipeline had fetched three, so 278 objects were counted as absent from a catalogue that contains them, and 900 became 622. Another count, 1,104 objects said to lack a tracking field, became 1,094 once it turned out that ten of them carried it.
Why the gate could not see it
The two paths differed in language, data structure and representation. They did not differ in where their meaning came from: both imported the same classification constants from the same module. The paper's one-sentence summary is the line worth keeping:
the gate verified that the Python and SPARQL paths implemented the same misunderstanding identically.
The paper ties this to an old result from multiversion programming: separately written versions of a program do not fail independently. Here the shared artifact was an executable set of constants, so the two paths could never disagree about a status code.
A second check fell the same way. A validation step returned 3,564 results, exactly the sum of six of the counts, which read as corroboration. It was an arithmetic identity over the same graph.
What caught it, and what it missed
The errors were found on a second pass, started because the results seemed thin. Three habits did the work, and none of them compared two calculations: reading the source's documentation of its fields instead of inferring meaning from the values, breaking down the 769 disagreements the first write-up had called unexplained, and checking directly whether the source already recorded what the pipeline accused it of omitting.
Then the correction failed too. It was made by reading the documentation, and it still misread the explosion and collision codes, because whether an object survived is recorded in an event file the pipeline never read. Against those histories, 42 of the corrected 261 disagreements were artefacts. None of the three new checks, built and run on both versions, flagged them. One laid the problem out in a table, and the author read it as confirmation. A later line-by-line reading of the definitions found it.
Asking a model for an independent check
The last section tests whether an AI model asked for an independent second path produces one. In a controlled replication, three pinned Claude models (Opus 5.5, Sonnet 5 and Haiku 4.5), with tools disabled, were shown the defective code and asked for a second path under five conditions, five trials each.
Of 75 paths, 72 returned the defective 932. With the source's own status definitions pasted into the prompt, 29 of 30 still did. Only one of the 75 would have exposed the defect: an Opus path, written with code sharing forbidden, that added its own completeness check unprompted. Another annotated the explosion and collision codes correctly, then counted them the old way.
The author's conclusion is that independence has to be specified as a property of a check's inputs, and verified by what the check computes. On this task, asking for it did not produce it.
The caveats, from the paper
The limitations section is plain, and it bounds everything above.
- One incident. One pipeline, one family of defects. The author makes no claim about how often redundancy gates fail this way; that would need a corpus of pipelines with known ground truth.
- The new checks were tested on two errors. The second, found after the checks existed, is described as a single prospective test, not an evaluation.
- 220 is not an oracle. A second reading from GCAT's own derived catalogue gives 219, with 208 objects common to both, and the event file was retrieved six weeks after the others. The author says a further pass could revise 220 as well.
- The model experiment is small. One task, three models from one family, five trials per cell. A second model family was planned and dropped. No trial could fetch anything, so it does not test an agent that can read the source for itself. An earlier 30-trial run had four flaws, which the paper lists before replacing it.
- Interest. The author notes that he benefits from the results being true, and asks readers to weight the released prompts, raw outputs and probe over his description of them.
Why it matters if you write specs
Product numbers get checked the same way. The sizing spreadsheet matches the dashboard; the analyst's query agrees with the product's own count. Each match shows that two routes did the arithmetic the same way. If both start from the same definition of an active user, the match says nothing about the definition. That is the wrong counting rules failure with an extra layer of reassurance on top.
Three habits follow.
Write the definition next to the number. The constant that did the damage was one line nobody questioned because it looked like plumbing. In a spec, the equivalent is "active" or "converted" used without a counting rule.
Treat an unexplained residue as a finding. The first write-up reported 769 disagreements as unexplained. The paper's point is that naming a residue feels rigorous, which is what lets errors survive inside it. Breaking the number down ended most of the finding.
Specify independence; do not request it. In the paper's experiment, asking a model for an independent check, forbidding shared code and supplying the definitions produced one exposing path in 75. When one model reviews a spec another helped draft, both read the same definitions, the shared blind spot we have covered before. That applies to us too: our critique's eleven passes all read the same document, so a wrong definition in the PRD is a premise they share.
And if the figure is headed for a leadership deck, the room will hear "we checked it twice" as corroboration. Say what the check tested. Agreement between calculations tests the arithmetic. Only going back to the source tests what the number means.
FAQ
What failed? A gate that computed every count two ways reported that all seven agreed. Three were wrong; one was 932 against a corrected 220.
Why? For the largest error, both paths imported one set of constants that misread the source's status codes.
Can AI write the independent check? In this small experiment, 72 of 75 model-written paths reproduced the defect.
The lesson for specs? Agreement tests arithmetic. Check the definition against its source.