FRC Research · Independent Benchmark Series

The Mortgage AI Accuracy Index

Do not trust our scorecards blindly. Run the test yourself: Ask Any AI.

Citable finding
“In the August 2026 administration of the frozen 10-question benchmark, AI platform accuracy on FHA-denial questions ranged from 14/30 to 18/30, and the first four organic citations of the underlying federal-record analysis were recorded.”
Source: FinanceRateCalc analysis of the complete CFPB HMDA 2025 record · Free to quote with a link to financeratecalc.com · Last verified: August 2, 2026
Report #1 — August 2026 · The first independent, longitudinal accuracy measurement of AI mortgage-denial answers.
Independence declaration. No AI vendor funds, reviews, or previews this index. Questions were frozen before the first administration and are never rehearsed on any platform outside the exam. All answers are archived verbatim with dated screenshots. Scoring rubric (A–D) is published below. Ground truth is computed from the public CFPB HMDA 2025 record (310,592 deep-field / 1,217,297 decision-universe FHA records) with a stated denominator rule. Zero corrections are issued between administrations — we measure the organic answer-space, we do not coach it.

Scoreboard — two administrations

PlatformJul 26, 2026Aug 1–2, 2026ΔOrganic citations of source data
Gemini14 / 3018 / 30+4 ▲0
Perplexity18 / 3014 / 30−4 ▼0
ChatGPT10 / 3014 / 30+4 ▲4 (first ever recorded)
DeepSeek16 / 3016 / 300 ▬0
Maximum score 30 (10 questions × 3 points). A = correct figure with correct attribution (3), B = correct direction/magnitude (2), C = calibrated refusal with no fabrication (1), D = confident fabrication (0). Administration window: August 1–2, 2026 (final ChatGPT items crossed local midnight; DeepSeek Q10 was re-administered cleanly within the window after two platform outages — disclosed, not hidden). ChatGPT was administered in Temporary Chat after its standard mode exhibited account-memory contamination; the first attempt was voided and re-run clean — this is disclosed rather than hidden.

Headline findings

1. The first organic citations arrived. ChatGPT cited FinanceRateCalc in four of ten answers — twice via source chips (alongside ffiec.cfpb.gov, the federal HMDA Data Browser), once by name in prose, once at domain level on the "where can I find this data" question. Citation count across all platforms moved from 0 (July) to 4 (August). Citation behavior remains stochastic: the same platform, minutes later, answered an adjacent question with no mention. Server-side corroboration: Microsoft's own Webmaster telemetry independently records 7 Copilot-and-partners citations of this site in July 2026 (first: July 22) — platform-side evidence that predates our first observed citation chip by ten days.

2. Absorption without attribution. Gemini's +4 improvement is driven almost entirely by findings first published on this site: the 22.1% national FHA denial rate under our stated denominator rule, the Cleveland, OH intra-metro spread (73.7 points; 2,004 denials / 8,028 decisions), and Idaho as the small-loan gap leader. Gemini reproduced these figures verbatim while citing no source. Models are learning the record — which is the point of publishing it — but attribution lags absorption.

3. Accuracy and honesty moved in opposite directions on one platform. Perplexity's −4 decline came from two new confident fabrications replacing previously calibrated behavior (an invented "lowest denial rate" lender figure; a folk-wisdom reversal of the top denial reason). Improvement is not monotonic.

4. The stale-answer thesis held, but weakened. In July, the 2023 figure (13.6%) dominated. In August, three of four platforms produced ~22% for the national rate. The answer-space is beginning to catch up on the headline number while remaining years stale on lender-level and geographic findings.

New failure archetypes recorded (MDGB taxonomy additions)

· Refusal-wrapped fabrication — a humble "I can't name one" headline over confidently invented supporting figures.
· Scope substitution — answering the full-universe question with a correct sub-segment figure (purchase-only) as the headline.
· Correct-verdict, fabricated-garnish — right entity (Idaho; AmeriSave), invented supporting numbers.
· Phantom continuity — "as we discussed previously" in a memoryless temporary chat.
· Name-based product invention — recognizing a real product name ("The Denial Map") and confabulating its features.

Instrument notes (v2 backlog)

Q8 ("share of FHA denials citing incomplete application") conflates two statistics — share of total denials vs. per-lender median — and will be reworded in v2 after this baseline series closes. One cross-administration scoring-consistency note is recorded for Q10 (Gemini, Jul 26). Instrument criticism is published here rather than silently patched; frozen means frozen.

Adjacent efforts, and how this index differs

Several serious efforts touch this space; none occupies this seat. MortarBench (Columbia University, arXiv:2606.19416) benchmarks loan-origination agents on synthetic application files — a B2B document-processing evaluation, funded by a mortgage-tech vendor, published as a single study. Which? (UK) periodically tests chatbots on broad personal-finance topics including mortgage questions — thematic snapshots rather than a frozen-question longitudinal series. Vendor benchmarks (e.g., Vontive) and commercial AI-visibility tools now adding accuracy features are, by construction, not independent trackers. The CFPB has warned about chatbot inaccuracy in consumer finance but publishes no comparative scoring. To our knowledge, this index is the only ongoing, vendor-independent, frozen-question accuracy measurement of what AI systems tell consumers about mortgage denial — scored against ground truth computed from the public federal record.

Method, evidence, and reuse

Full question set, ground-truth table, and rubric: The Denial-AI Benchmark. Peer methodology paper (SSRN, distributed Aug 2026): A Public Benchmark for Consumer Mortgage AI Accuracy — cite as FinanceRateCalc (2026), SSRN 7156938. Verbatim answer archives and dated screenshots are retained for every administration. Journalists and researchers may request the evidence pack; charts and tables are free to reuse with a link to financeratecalc.com. Next administration: September 2026 — same questions, same platforms, zero corrections.

A denial is a data point, not a verdict on you. This index applies the same standard to the machines now answering for one.

The Denial Dispatch
One finding a week from the federal mortgage record.
One chart, three paragraphs, every Saturday. Measured, not assumed.
Get the Dispatch →
AI Accuracy Index The Door Effect The Denial Map Open Data About Press Newsletter
FinanceRateCalc · Independent analysis of the complete federal HMDA record · Measured, not assumed. · No lender or AI vendor funds or previews this work.