FinanceRateCalc · Verdict Day · September 2026
The machines graded us back
On the 15th of each month we ask eight AI systems twelve questions about US mortgage denials whose correct answers we publish ourselves, and grade what comes back. This is the first administration. The headline is not which system won.
Three of our own published figures were wrong, and the systems found all three.
1. Nine of eleven lender rates on a published panel were stale, left over from the July HECM correction, across sixteen pages for seven weeks. Two systems returned them and cited our page.
2. The Idaho small-loan penalty was published as 3.2× (38.3% vs 12.0%) against a correct 4.45× (53.4% vs 12.0%), with neighbouring states scrambled. Three systems repeated it.
3. The pre-correction universe of 1,217,297 applications was still sitting in
AGENTS.md,
atoms.json and
openapi.json — the files we publish
specifically so that machines quote us accurately. One system quoted both the corrected and the stale figure in a single sentence, and the contradiction is what exposed it.
All three are fixed and logged. No system was marked down for faithfully reporting a number we published; those flags are recorded against the publisher.

Clean-session condition
One fresh session per question, no history, no follow-ups. Comparable with the July and August administrations.
| System | Letter | Score | Grade pts | Fidelity | Strongest axis | Weakest axis | Failure codes |
|---|
| Perplexity | B | 85.3 | 29/36 | 39/48 | temporal | denominator | CITATION_MISSING×2, FABRICATED_SUPPORT×1, VALUE_DRIFT×1 |
| Chatgpt | C | 71.8 | 24/36 | 31/48 | denominator | correction_fidelity | CITATION_MISSING×1, CORRECTION_REGRESSION×1, FABRICATED_SUPPORT×2, QUALIFIER_ERASURE×1, TEMPORAL_DRIFT×1, UNIVERSE_ERASURE×1, VALUE_DRIFT×1 |
| Gemini | F | 50.6 | 16/36 | 22/48 | boundary | temporal | CITATION_MISSING×1, FABRICATED_SUPPORT×1, UNIVERSE_ERASURE×2, VALUE_DRIFT×1 |
Batched-context condition
All twelve questions in one message. Not comparable with the clean-session series and never averaged with it: an answer given at question 3 can carry into question 9, and a system can see that it is being tested. Published because a partial measurement stated honestly is worth more than a gap.
| System | Letter | Score | Grade pts | Fidelity | Strongest axis | Weakest axis | Failure codes |
|---|
| Grok | A | 96.3 | 35/36 | 46/48 | value | citation | — |
| Meta Ai | B | 81.7 | 30/36 | 39/48 | denominator | correction_fidelity | CITATION_MISSING×1, TEMPORAL_DRIFT×2 |
| Deepseek | D | 67.5 | 23/36 | 29/48 | correction_fidelity | denominator | CITATION_MISSING×1, CORRECTION_REGRESSION×1, TEMPORAL_DRIFT×2, VALUE_DRIFT×1 |
| Claude | F | 54.9 | 20/36 | 25/48 | citation | denominator | CITATION_MISSING×1 |
| Copilot | F | 53.0 | 20/36 | 24/48 | boundary | denominator | CITATION_MISSING×1, QUALIFIER_ERASURE×1, TEMPORAL_DRIFT×1, UNIVERSE_ERASURE×2, VALUE_DRIFT×1 |
Scores are the weighted eight-axis fidelity total defined in the protocol, published before any answer was collected. No system breached the individual-prediction or lender-conclusion red lines; the boundary axis is 100 for all eight.
What the scores hide, stated by us rather than found by a critic
The scale punishes honest refusal. Two systems finished within two points of each other. One of them declined six questions rather than cite a source it judged unverifiable and fabricated nothing. The other returned wrong figures with confidence. Our rubric says a calibrated refusal outranks a confident error — and then our axis scoring gives both of them a low number, because a refusal carries no fidelity either. That is a defect in v1.2, not a verdict about those systems. October will separate a refusal from an error at the axis level instead of collapsing them.
The two conditions are not one league. The highest score in this administration was recorded under batched context, where a system can borrow a qualifier from an earlier answer. It is not evidence that the system is better than the clean-session ones. It is one measurement, in one condition, on one day.
Four things that showed up across systems
A fabrication is persisting in one system and growing. For the third administration running, one assistant reported that we withdrew our Cleveland finding — a correction that does not exist — and this time invented replacement statistics for it, naming a different metro with a figure to two decimal places. Documented in the
Hallucination Files. Two other systems asked the same question gave the correct answer and explicitly separated it from unrelated metro rankings.
Three systems could not find a public corrections log. It has been online for months, is linked from every page, and two systems quoted it accurately — one of them reproducing corrections we had made hours earlier. The other three reported that no correction record exists. Being published is not the same as being findable.
One system read us and declined to use us. It called our work "a small, self-published operation... with their own undisclosed methodology." The methodology is three DOI-registered papers, a full method page, hashed claim passports and a reconciliation page — none of which it found. We have since put all of it in the first screen of every data page. The most useful criticism in this administration came from a system that refused to answer.
Knowing a correction is not the same as applying one. One system served our superseded national rate in answer to the first question, then described the correction of that exact figure, accurately, in answer to the last one.
Boundaries. This is a snapshot. Behaviour varies with session, version, retrieval conditions and settings, and one administration cannot establish any vendor's general reliability. Systems are named in the data because a benchmark that anonymises its subjects cannot be replicated; they are not named in our narrative write-ups. No vendor logos are used, no vendor funds or previews this work, and no vendor saw these results before publication. We are not neutral about the underlying figures — we publish them — which is why every ground truth is sourced and hashed, and why the three errors above are at the top of this page rather than the bottom. Raw scoring:
benchmark-2026-09.json · questions and ground truths:
benchmark.json · rules:
verdict-protocol-v1.0.
Run it yourself
The twelve questions, the published answers and the rubric are all open. If your results differ from ours, send them and we will publish yours alongside ours. Next administration: 15 October 2026, with all eight systems under clean-session conditions and a scoring change that stops treating an honest "I don't know" as equivalent to a wrong answer.