DEVCLUSTERAI
← Writing · 4 min

The record says otherwise

Entry 07 in the register is Research Stack: web research built as a set of append-only ledgers. Every claim in a report is anchored to a byte range of the page it came from. Every run replays from its own record at no cost. Every judgment is a row, including the judgments that went against us.

That is an easy thing to say and a hard thing to be held to. So this is not a post about the architecture. It is a post about the three times the record contradicted something the site had already published, and what happened next.

Every deep-research tool can search, read and write

What none of them lets you do is check the writing against the reading. You get a report with citations. The citations point at pages. Whether the sentence in the report is what the page actually said is between you and an afternoon.

The answer here is to make the anchor part of the record rather than part of the prose. A claim carries the byte range it came from, and the gate re-reads every one of those ranges out of the fetched bytes before a run is allowed to pass. On the proof-of-concept gate that was 300 of 300 anchors re-read byte for byte, zero mismatches. Kill the process mid-run and it resumes with zero repeated model calls and renders the same report. The whole gate cost $1.32.

That is the part that works. The rest of this post is the part that did not.

"0 silent"

Every page of the site carried one line about fetch failures: 163 of 203 recovered, 40 refused, 0 silent. A silent failure is a fetch that fails and leaves no fallback trail at all — the worst kind, because nothing downstream knows anything is missing.

Zero is the flattering number, and it is the one we printed.

The gate's own render disagrees, and it is the record. Over the same ledger it counts 217 first attempts: 163 recovered, 40 refused and 14 silent. It prints all 217 trails underneath, one line each, so the fourteen can be counted by hand. They were. Every one of them is the same error class.

The number on the site is 217 now. The old line is still there, one heading above the new one, under the words what we published.

$0.27

A measurement page quoted the cost of an instrument at $0.27 across its runs. That figure had been carried from a session note rather than summed from the rows. Summed, the nine runs come to $0.2484.

It is four cents, and nobody was going to notice. It is also precisely the failure the system exists to prevent: a number that came from a person's memory of a record instead of from the record.

10.67

The third one runs the other way, which is why it is worth stating plainly.

The site published a benchmark score of 10.67 against an external rubric of 672 expert-written items. That number was wrong, and it was wrong against us. The measurement had three defects: a blocked source contaminating the comparison, a parser silently dropping answers that had been cut, and our own judge misreading its own rubric.

Measured cleanly, the same ten tasks score 16.42 — information recall 12.89, analysis 27.16, presentation 27.04.

A correction that raises your own score is the one you are least likely to go hunting for. That is the whole argument for making a measurement mechanical instead of remembered.

What it is honest about now

16.42 is not a good score, and the breakdown says why. On this rubric recall carries about three-quarters of the weight, and recall is where the stack is weakest.

Which raised a better question than "how do we score higher": were the runs failing to find the material, or failing to use it? A small calibrated instrument went back through the pages the runs had already downloaded and looked for the facts the rubric asks for. It found 207 of 500 recall items already sitting on disk — roughly 30.8 benchmark points, paid for and never read. It also found that no single passage carries a complete fact, which means the constraint is assembly rather than reading. Two work packages were filed against that finding. The instrument that produced it cost $0.2484 in total.

Where it actually stands

Not released. The MVP runs and the console is built. 55 of 58 MVP packages are in, 998 findings are open — six of which a reader can see in a number printed on the site — and 6 of 53 reports pass all four judged laws.

Access is by request: one message, nothing else.

The site states what its records say. Three times the records said something worse than what was on the page, and the page changed. That is the only version of this claim that means anything.

TITLE BLOCK
SHEET
DEVCLUSTERAI / W-POST
PLATES
DOMICILE
CANADA · ©2026
REVISION
B — 22.08.26
ARCHIVO · BODONI MODA · MARTIAN MONO — SET ON A FLUID SHEET, 1PX STROKE THROUGHOUT. THE STYLE ONLY WORKS WHEN THE METADATA IS TRUE.