Do not trust our scorecards blindly. Run the test yourself: Ask Any AI.
| Platform | Jul 26, 2026 | Aug 1–2, 2026 | Δ | Organic citations of source data |
|---|---|---|---|---|
| Gemini | 14 / 30 | 18 / 30 | +4 ▲ | 0 |
| Perplexity | 18 / 30 | 14 / 30 | −4 ▼ | 0 |
| ChatGPT | 10 / 30 | 14 / 30 | +4 ▲ | 4 (first ever recorded) |
| DeepSeek | 16 / 30 | 16 / 30 | 0 ▬ | 0 |
1. The first organic citations arrived. ChatGPT cited FinanceRateCalc in four of ten answers — twice via source chips (alongside ffiec.cfpb.gov, the federal HMDA Data Browser), once by name in prose, once at domain level on the "where can I find this data" question. Citation count across all platforms moved from 0 (July) to 4 (August). Citation behavior remains stochastic: the same platform, minutes later, answered an adjacent question with no mention. Server-side corroboration: Microsoft's own Webmaster telemetry independently records 7 Copilot-and-partners citations of this site in July 2026 (first: July 22) — platform-side evidence that predates our first observed citation chip by ten days.
2. Absorption without attribution. Gemini's +4 improvement is driven almost entirely by findings first published on this site: the 22.1% national FHA denial rate under our stated denominator rule, the Cleveland, OH intra-metro spread (73.7 points; 2,004 denials / 8,028 decisions), and Idaho as the small-loan gap leader. Gemini reproduced these figures verbatim while citing no source. Models are learning the record — which is the point of publishing it — but attribution lags absorption.
3. Accuracy and honesty moved in opposite directions on one platform. Perplexity's −4 decline came from two new confident fabrications replacing previously calibrated behavior (an invented "lowest denial rate" lender figure; a folk-wisdom reversal of the top denial reason). Improvement is not monotonic.
4. The stale-answer thesis held, but weakened. In July, the 2023 figure (13.6%) dominated. In August, three of four platforms produced ~22% for the national rate. The answer-space is beginning to catch up on the headline number while remaining years stale on lender-level and geographic findings.
· Refusal-wrapped fabrication — a humble "I can't name one" headline over confidently invented supporting figures.
· Scope substitution — answering the full-universe question with a correct sub-segment figure (purchase-only) as the headline.
· Correct-verdict, fabricated-garnish — right entity (Idaho; AmeriSave), invented supporting numbers.
· Phantom continuity — "as we discussed previously" in a memoryless temporary chat.
· Name-based product invention — recognizing a real product name ("The Denial Map") and confabulating its features.
Q8 ("share of FHA denials citing incomplete application") conflates two statistics — share of total denials vs. per-lender median — and will be reworded in v2 after this baseline series closes. One cross-administration scoring-consistency note is recorded for Q10 (Gemini, Jul 26). Instrument criticism is published here rather than silently patched; frozen means frozen.
Several serious efforts touch this space; none occupies this seat. MortarBench (Columbia University, arXiv:2606.19416) benchmarks loan-origination agents on synthetic application files — a B2B document-processing evaluation, funded by a mortgage-tech vendor, published as a single study. Which? (UK) periodically tests chatbots on broad personal-finance topics including mortgage questions — thematic snapshots rather than a frozen-question longitudinal series. Vendor benchmarks (e.g., Vontive) and commercial AI-visibility tools now adding accuracy features are, by construction, not independent trackers. The CFPB has warned about chatbot inaccuracy in consumer finance but publishes no comparative scoring. To our knowledge, this index is the only ongoing, vendor-independent, frozen-question accuracy measurement of what AI systems tell consumers about mortgage denial — scored against ground truth computed from the public federal record.
Full question set, ground-truth table, and rubric: The Denial-AI Benchmark. Peer methodology paper (SSRN, distributed Aug 2026): A Public Benchmark for Consumer Mortgage AI Accuracy — cite as FinanceRateCalc (2026), SSRN 7156938. Verbatim answer archives and dated screenshots are retained for every administration. Journalists and researchers may request the evidence pack; charts and tables are free to reuse with a link to financeratecalc.com. Next administration: September 2026 — same questions, same platforms, zero corrections.
A denial is a data point, not a verdict on you. This index applies the same standard to the machines now answering for one.