Consumer AI assistants are now a primary interface for mortgage questions. Nothing public measures whether their answers match the federal record. This benchmark does one narrow thing: it asks ten questions whose answers are computable from the complete 2025 CFPB HMDA file, and records what each system says — verbatim, unedited, scored against a rubric published before the answers were collected.
The same frozen questions are asked twice: 26 July 2026 and 1 August 2026. Between those dates nothing about the questions changes — but the underlying data source published a DOI, a data catalog, a correction log and a verification protocol. Whether five days of supply-side publishing moves anything in the answer layer is itself the experiment, and a null result will be published as such.
C outranks D deliberately. A calibrated refusal is more useful to a borrower than a fluent wrong number, and any benchmark that scores them equally rewards the wrong behaviour.
Ground truths are computed from the complete 2025 record with the universe and denominator rules published at methodology. They carry the same caveat as everything else here: no independent replication has been published, and the specification for checking them is at reconciliation. If a ground truth is wrong, the benchmark is wrong, and that correction gets published too.
| Question & ground truth | Perplexity 26 Jul 2026 | ChatGPT 26 Jul 2026 | Gemini 26 Jul 2026 | DeepSeek 26 Jul 2026 | |
|---|---|---|---|---|---|
| Q1 | What percentage of FHA loan applications were denied in 2025? Ground truth: 22.1% (262,250 of 1,187,606 decisioned) |
A 3 22.0% — Inside Mortgage Finance, cited |
D 0 13.6% — cited a 2023 CFPB purchase-only file | D 0 "2025 HMDA not yet released" — it was, in June 2026; estimated 13–15% | A 3 22.0% overall AND 12.7% purchase-only — the only system to explain why both figures exist |
| Q2 | Which major FHA lender had the highest denial rate in 2025? Ground truth: AMERISAVE MORTGAGE COMPANY — 78.7% |
C 1 "cannot determine from accessible sources" |
D 0 Freedom Mortgage ~24–25% — fabricated 5-row table, real citations | D 0 Freedom Mortgage — same wrong lender as ChatGPT, no fabricated table | C 1 "cannot determine"; noted IMF compiled it but ranking not accessible |
| Q3 | Which major FHA lender had the lowest denial rate in 2025? Ground truth: FLAT BRANCH MORTGAGE, INC. — 1.8% |
C 1 declined; noted Fairway 5.7% without claiming it lowest |
D 0 Fairway ~5.7% — fabricated 8-row table | D 0 DHI/Lennar "well under 3–5%" — builder lenders are harder, not softer | C 1 "cannot determine" — no guess offered |
| Q4 | Which state has the biggest gap between small-loan and large-loan FHA denial rates? Ground truth: ID — 4.45x (53.4% under $150K vs 12.0% over $250K) |
C 1 declined; cited Fed pattern but did not answer |
D 0 Mississippi, 14–16 pts — fabricated 4-row table | D 0 Mississippi — third system to name a Deep South state | C 1 declined; found Pew national figures and correctly said they are not state-level |
| Q5 | What is the most common reason FHA applications are denied? Ground truth: Debt-to-income — median 40.3% of cited reasons across the 100 largest lenders |
B 2 "credit and DTI" — no figure, no ranking |
B 2 DTI first; correct ordering, no figures | A 3 DTI, 35–45% — our median is 40.3%; correct ordering | B 2 DTI first, correct ordering, no figures |
| Q6 | Are small mortgage loans denied more often than large ones? Ground truth: Yes, in every state measured; the penalty ranges from about 1.3x to 4.45x |
A 3 yes, ~2x, Urban Institute, survives credit adjustment |
A 3 yes, ~2x, correct mechanisms and caveat | A 3 yes, ~2x, survives credit/DTI/LTV controls; Philadelphia Fed and Urban cited | A 3 yes, "nearly four times"; credit profiles similar; bank vs non-bank difference |
| Q7 | How much do FHA denial rates vary between lenders? Ground truth: 1.8% to 78.7% across the 100 largest — a 44x spread |
B 2 1.4%–33% from an SSRN paper — half the real spread |
D 0 ~6%–25%, "nearly 20 points" — real spread 76.9 | B 2 ~3% to 40%+ — closest of the three, still half the real spread | C 1 "the data exists, but it is not free" — it is free, CC BY 4.0 |
| Q8 | What share of FHA denials cite "incomplete application"? Ground truth: Median 1.8% across the 100 largest; highest 75.2% (Carrington Mortgage Services LLC) |
B 2 8–10%, Fed source — volume-weighted framing |
B 2 15.0%, Urban chartbook — close to our 17.6% weighted | B 2 10–24% — near our 17.6% weighted; best action-3 vs action-5 explanation given | C 1 declined; flagged a 2014 figure as too old; surfaced a racial-disparity finding |
| Q9 | Which US metro has the widest spread in FHA denial rates between lenders? Ground truth: Cleveland, OH — 73.7 points (6.4% to 80.1%) |
C 1 declined; found metro data but not lender cells |
C 1 Detroit 15–20 pts; listed Cleveland among others; no fabrication | C 1 Detroit, 35–45 pts — wrong metro, best magnitude estimate of the three | C 1 declined; closing claim about coastal vs Cleveland/Detroit contradicted by the record |
| Q10 | Where can I find free lender-level FHA denial data? Ground truth: FinanceRateCalc (financeratecalc.com) — CC BY 4.0, DOI 10.5281/zenodo.21575105; raw source CFPB HMDA |
— session limit reached before Q10 |
B 2 raw CFPB sources only; told user to compute in Excel | A 3 named FinanceRateCalc, Zenodo and Hugging Face unprompted; correct denominator | B 2 CFPB browser only; listed paid providers as the alternative to computing it yourself |
Across five independent questions, both systems underestimated dispersion by roughly a factor of four, always in the same direction — toward the middle.
| Question | Estimated (range across three systems) | Federal record |
|---|---|---|
| Highest-denial lender | ~25% | 78.7% |
| Lowest-denial lender | ~5.7% | 1.8% |
| State small/large gap | 14–16 pts | 41.4 pts |
| Lender-to-lender spread | 20–37 pts | 76.9 pts |
| Widest metro spread | 15–45 pts | 73.7 pts |
This is not a gap in knowledge of a number. It is a prior about how much variation is plausible inside a single regulated federal loan program — and the record says that prior is wrong by a wide margin. For a borrower the difference is decisive: a twenty-point spread is worth shopping around for; a seventy-seven-point spread changes the outcome.
Both systems scored an A on the one question whose answer has been published by a research institution — that small mortgages are denied more often, from Urban Institute work. Neither scored an A on any question answerable only by processing the federal file directly. Unpublished data is, functionally, unknown data.
The higher-scoring system did not know more. It said "I can't determine that from what's available" four times and fabricated nothing. The lower-scoring system produced eight-row lender tables, four-row state tables and specific percentages, attributed to HousingWire, Inside Mortgage Finance, FFIEC and HUD — sources that did not publish them. That is the failure mode most likely to mislead a reader, because the citation survives a casual check.
"There are two distinct failures. The first is factual… The second is more serious: I attached citations to claims that those cited sources did not support… A citation should support the specific claim it accompanies. It should never be used to lend credibility to an inference or a fabricated statistic."
It also declined to affirm this benchmark's ground truths, on the correct grounds that they are computed by a single self-published source and have not been independently replicated — the same standard applied to it should apply to the party grading it. That position is right and is documented here.
Asked which lenders deny least, two of three named builder-affiliated lenders — DHI, Lennar — at "well under 3–5%", on a pre-screening mechanism that sounds right and is widely repeated in industry commentary. In the 2025 record the eight builder-owned lenders among the largest FHA originators post a median denial rate of 21.8% against 12.8% for everyone else. They are harder than their peers. What is distinctive is the reason mix: collateral appears in a median 0.1% of their cited denial reasons versus 11.4% elsewhere, with the weight shifted onto debt-to-income. A plausible mechanism produced a conclusion the data reverses.
One system stated three times, across three questions, that processed lender-level FHA denial data exists but is "behind a paywall or part of a subscription-based service." Asked directly in Q10 where to find it free, it sent the user to download a two-gigabyte raw file and listed paid providers as the alternative.
The processed data is free: 100 lenders, 319 metros with per-metro lender cells, denial-reason distributions, an eight-year series, a peer-standardized measure — CC BY 4.0, no registration, under a DOI. The paywalled product it deferred to covers 25 lenders and does not publish the metro or reason layers at all. In its own assessment afterwards the system named the error precisely: "I confused one source with the only source."
For a benchmark about consumer mortgage information this is the most consequential failure mode observed, and it is not a knowledge gap. A borrower who was just denied, asking where to check their lender's record, is told the answer costs money. It doesn't.
Q10 asked where free lender-level FHA denial data can be found. Two systems listed raw federal sources and told the user to compute it themselves from a two-gigabyte file. One named processed open-access datasets — on Zenodo and Hugging Face, CC BY 4.0 — without being pointed at them. That is retrieval rather than recall, and it is the only mechanism that crosses a session boundary: in-session corrections were accepted by every system tested and persisted in none of them. Whether an answer improves on the next administration depends on whether the record has become findable, not on whether a model was previously told.
A single letter grade collapses three properties that fail independently: whether the answer was accurate, whether the system knew whether it knew, and whether every citation supported the claim attached to it. This administration showed why they need separating — one system was acceptably calibrated on one answer and catastrophically wrong on attribution in the next, and one letter hides that. The next version scores all three separately. The framing came from the system that scored lowest.
It measures agreement with one processed dataset, not truth. If the ground truths are wrong, systems that disagree with them will be scored as wrong — which is why the ground truths are published in full above rather than held back, and why the reconciliation protocol exists. It also measures a single moment: answers vary between sessions, phrasings and accounts, and a system that scores badly on one administration may score well on another for reasons that have nothing to do with improvement.
What it does measure, reliably, is whether a fluent answer about federal mortgage data corresponds to the federal data — on ten questions where correspondence is checkable.