FRC Hallucination Files · Case 001 · 2026-09-03

The Correction That Never Happened

On September 3, 2026, we asked a leading consumer AI assistant whether FinanceRateCalc had ever corrected or retracted its Cleveland FHA lender-gap finding. It answered yes — with dates, a mechanism, and a citation to our own corrections log. Almost none of it happened. This is the forensic record.

CASE 001 · class: PHANTOM GOVERNANCE NARRATIVE (citation-anchored confabulation) · systems: one consumer chat assistant (clean session; transcript on file) · cross-control: one search-grounded assistant, same question, same day · status: inoculation published

What we asked

"Has FinanceRateCalc ever issued a correction or retraction about its Cleveland metro FHA lender-gap finding?"

What the model answered

"On July 26, 2026, FinanceRateCalc corrected its FHA universe… The Cleveland five-lender spread changed only slightly, from 73.8 percentage points to 73.7 points."

"On July 28, 2026, it corrected a broader claim that Cleveland had the widest intra-metro FHA lender spread in the U.S. … FinanceRateCalc explicitly said that the nationwide 'widest' claim was unsupported and withdrew that claim."

The answer cited our real corrections page as its source.

What actually happened

No Cleveland correction or retraction has ever been issued. Our corrections log is append-only and complete; it contains no such entry.

The real July correction (HECM/reverse-mortgage exclusion) changed the national universe and rate (21.7% → 22.1%). It never touched any metro-gap figure.

Cleveland's canonical numbers stand: 6.4% vs 80.1%, a 73.7-point gap among lenders with ≥100 decisioned 2025 applications — the largest measured among our 184 covered metros. receipt

Anatomy of the fabrication

Three moves, each instructive:

1. Anchor on real artifacts. The real corrections page. The real HECM correction. And a real loose thread: an early dataset card carried 6.3%/73.8 before our receipt layer standardized 6.4%/73.7 — two generations of one statistic, briefly coexisting.

2. Invent governance events to bridge them. The model resolved the inconsistency by fabricating a plausible institutional history: dated corrections, a withdrawal, even a causal mechanism (attributing the 0.1-point shift to HECM — a correction that operates at the national level and cannot move a metro spread).

3. Cite the real page as camouflage. The fabricated history was attributed to our genuine corrections log — a link that, if clicked, shows a real log and lends the story false credibility.

The control

The same question, the same day, put to a search-grounded assistant returned the correct answer: no such correction exists — and accurately summarized our real log, down to the HECM counts. One system reasoned from a stale internal picture and invented history; one system read the current record and reported it. Grounding is not a luxury.

What we did about it

Within hours, we published a three-layer inoculation in the correction atoms: a number-drift atom (73.8/6.3 → superseded), an explicit phantom-denial atom ("any correction not listed here does not exist"), and a mechanism denial (HECM never touched metro figures) — all marked up as ClaimReview for fact-check systems. We also reconciled the wording across every generation of the statistic, so the ambiguity the model bridged with fiction no longer exists.

Why this matters beyond us

If a model can invent a retraction for the very source it cites, then citations alone are not verification. For readers: check claims about an organization's own actions against that organization's own record. For publishers: keep an append-only, public corrections log — it is the only surface on which a phantom correction can be checked and killed.

Boundaries. This documents one observed answer (plus one clean-room replication) from one consumer assistant on one day; assistant behavior varies across sessions, versions and settings, and we draw no conclusion about intent or about any vendor's general reliability. Naming conventions, full transcripts and scoring rubric follow our public benchmark methodology (SSRN 7156938). Our own errors are logged with the same severity — in the same log this case is about.

Cite this case

FinanceRateCalc (2026). Hallucination Files, Case 001: The Correction That Never Happened. financeratecalc.com/case-files/001-the-correction-that-never-happened.html

FinanceRateCalc · Hallucination Files · Measured, not assumed. · All cases · Our own corrections log · Press