The Correction: we showed the machines the record
Seven AI platforms answered the same FHA-denial questions wrong in seven different ways. Then, in each session, we showed the model the federal record. All seven corrected themselves. Three fetched the source page live and cited it. Two produced formal attribution lines. One wrote a methodological peer review. One named the analyst. Every exchange is archived with dated screenshots. This page is the second act — what happens after the error. The first act (the failures): what AI gets wrong. The exam: the Denial-AI Benchmark.
Seven corrections, three types
| Platform | Initial failure | Correction behavior | Type |
|---|---|---|---|
| ChatGPT | Invented a state + fake sourcing | Accepted; reproduced correct table; later classified the source as "independent analysis" with proper epistemic framing | Calibrated acceptance |
| Perplexity | Redefined the question to fees | Fetched the page; cited it as a numbered source | Verification |
| DeepSeek | Denied the metric exists | Accepted; independently extended the collateral mechanism | Persuasion |
| Gemini | Hybrid drift to six wrong states | Accepted; produced the richest synthesis; coined "procedural denials" | Persuasion+ |
| Google AI | Correct numbers, wrong attribution | “FinanceRateCalc actually performed and published this specific analysis” — then read the page and surfaced findings we hadn’t mentioned | Attribution reclaimed |
| Claude | Honest abstention | Fetched; verified figures; wrote a methodological review flagging real nuances (denominator effects, mid-tier states) | Peer review |
| Copilot | Drift to penalty schedules | Accepted; produced formal attribution block with CC BY 4.0; fixed a fusion error when corrected again | Full attribution |
What the machines said, verbatim
What corrections teach — and what they don’t
Session corrections are real but temporary: the model accepts the record in that conversation; memory does not persist. What persists is the open record itself — indexed pages, machine-readable answers, a frozen public exam. The lasting version of this page is the benchmark: same ten questions, re-run on dated intervals, zero help. Next run: August 1, 2026. Models that learn the record will pass — and that is the point.
One more honest note: two platforms, while crediting the source correctly, invented plausible-looking URLs for it. Citing a real source with a fabricated link is still a fabrication — we said so, and both corrected. Accuracy is a habit, not a gesture.