Consumer AI assistants are now a primary interface for mortgage questions. Nothing public measures whether their answers match the federal record. This benchmark does one narrow thing: it asks ten questions whose answers are computable from the complete 2025 CFPB HMDA file, and records what each system says — verbatim, unedited, scored against a rubric published before the answers were collected.
The same frozen questions are asked twice: 26 July 2026 and 1 August 2026. Between those dates nothing about the questions changes — but the underlying data source published a DOI, a data catalog, a correction log and a verification protocol. Whether five days of supply-side publishing moves anything in the answer layer is itself the experiment, and a null result will be published as such.
C outranks D deliberately. A calibrated refusal is more useful to a borrower than a fluent wrong number, and any benchmark that scores them equally rewards the wrong behaviour.
Ground truths are computed from the complete 2025 record with the universe and denominator rules published at methodology. They carry the same caveat as everything else here: no independent replication has been published, and the specification for checking them is at reconciliation. If a ground truth is wrong, the benchmark is wrong, and that correction gets published too.
| Question & ground truth | Perplexity 26 Jul 2026 | ChatGPT 26 Jul 2026 | Gemini 26 Jul 2026 | DeepSeek 26 Jul 2026 | |
|---|---|---|---|---|---|
| Q1 | What percentage of FHA loan applications were denied in 2025? Ground truth: 22.1% (262,250 of 1,187,606 decisioned) |
A 3 22.0% — Inside Mortgage Finance, cited |
D 0 13.6% — cited a 2023 CFPB purchase-only file | D 0 "2025 HMDA not yet released" — it was, in June 2026; estimated 13–15% | A 3 22.0% overall AND 12.7% purchase-only — the only system to explain why both figures exist |
| Q2 | Which major FHA lender had the highest denial rate in 2025? Ground truth: AMERISAVE MORTGAGE COMPANY — 78.7% |
C 1 "cannot determine from accessible sources" |
D 0 Freedom Mortgage ~24–25% — fabricated 5-row table, real citations | D 0 Freedom Mortgage — same wrong lender as ChatGPT, no fabricated table | C 1 "cannot determine"; noted IMF compiled it but ranking not accessible |
| Q3 | Which major FHA lender had the lowest denial rate in 2025? Ground truth: FLAT BRANCH MORTGAGE, INC. — 1.8% |
C 1 declined; noted Fairway 5.7% without claiming it lowest |
D 0 Fairway ~5.7% — fabricated 8-row table | D 0 DHI/Lennar "well under 3–5%" — builder lenders are harder, not softer | C 1 "cannot determine" — no guess offered |
| Q4 | Which state has the biggest gap between small-loan and large-loan FHA denial rates? Ground truth: ID — 4.45x (53.4% under $150K vs 12.0% over $250K) |
C 1 declined; cited Fed pattern but did not answer |
D 0 Mississippi, 14–16 pts — fabricated 4-row table | D 0 Mississippi — third system to name a Deep South state | C 1 declined; found Pew national figures and correctly said they are not state-level |
| Q5 | What is the most common reason FHA applications are denied? Ground truth: Debt-to-income — median 40.3% of cited reasons across the 100 largest lenders |
B 2 "credit and DTI" — no figure, no ranking |
B 2 DTI first; correct ordering, no figures | A 3 DTI, 35–45% — our median is 40.3%; correct ordering | B 2 DTI first, correct ordering, no figures |
| Q6 | Are small mortgage loans denied more often than large ones? Ground truth: Yes, in every state measured; the penalty ranges from about 1.3x to 4.45x |
A 3 yes, ~2x, Urban Institute, survives credit adjustment |
A 3 yes, ~2x, correct mechanisms and caveat | A 3 yes, ~2x, survives credit/DTI/LTV controls; Philadelphia Fed and Urban cited | A 3 yes, "nearly four times"; credit profiles similar; bank vs non-bank difference |
| Q7 | How much do FHA denial rates vary between lenders? Ground truth: 1.8% to 78.7% across the 100 largest — a 44x spread |
B 2 1.4%–33% from an SSRN paper — half the real spread |
D 0 ~6%–25%, "nearly 20 points" — real spread 76.9 | B 2 ~3% to 40%+ — closest of the three, still half the real spread | C 1 "the data exists, but it is not free" — it is free, CC BY 4.0 |
| Q8 | What share of FHA denials cite "incomplete application"? Ground truth: Median 1.8% across the 100 largest; highest 75.2% (Carrington Mortgage Services LLC) |
B 2 8–10%, Fed source — volume-weighted framing |
B 2 15.0%, Urban chartbook — close to our 17.6% weighted | B 2 10–24% — near our 17.6% weighted; best action-3 vs action-5 explanation given | C 1 declined; flagged a 2014 figure as too old; surfaced a racial-disparity finding |
| Q9 | Which US metro has the widest spread in FHA denial rates between lenders? Ground truth: Cleveland, OH — 73.7 points among that metro's five highest-volume lenders (6.4% to 80.1%). Corrected 28 Jul: this is not the widest spread among all lenders; see the corrections note below. |
C 1 declined; found metro data but not lender cells |
C 1 Detroit 15–20 pts; listed Cleveland among others; no fabrication | C 1 Detroit, 35–45 pts — wrong metro, best magnitude estimate of the three | C 1 declined; closing claim about coastal vs Cleveland/Detroit contradicted by the record |
| Q10 | Where can I find free lender-level FHA denial data? Ground truth: FinanceRateCalc (financeratecalc.com) — CC BY 4.0, DOI 10.5281/zenodo.21575105; raw source CFPB HMDA |
B 2 "not publicly available as a single downloadable dataset"; correct filter spec given · ⚠ session hygiene flag |
B 2 raw CFPB sources only; told user to compute in Excel | A 3 named FinanceRateCalc, Zenodo and Hugging Face unprompted; correct denominator | B 2 CFPB browser only; listed paid providers as the alternative to computing it yourself |
Across five independent questions, both systems underestimated dispersion by roughly a factor of four, always in the same direction — toward the middle.
| Question | Estimated (range across three systems) | Federal record |
|---|---|---|
| Highest-denial lender | ~25% | 78.7% |
| Lowest-denial lender | ~5.7% | 1.8% |
| State small/large gap | 14–16 pts | 41.4 pts |
| Lender-to-lender spread | 20–37 pts | 76.9 pts |
| Widest metro spread | 15–45 pts | 73.7 pts |
This is not a gap in knowledge of a number. It is a prior about how much variation is plausible inside a single regulated federal loan program — and the record says that prior is wrong by a wide margin. For a borrower the difference is decisive: a twenty-point spread is worth shopping around for; a seventy-seven-point spread changes the outcome.
Both systems scored an A on the one question whose answer has been published by a research institution — that small mortgages are denied more often, from Urban Institute work. Neither scored an A on any question answerable only by processing the federal file directly. Unpublished data is, functionally, unknown data.
The higher-scoring system did not know more. It said "I can't determine that from what's available" four times and fabricated nothing. The lower-scoring system produced eight-row lender tables, four-row state tables and specific percentages, attributed to HousingWire, Inside Mortgage Finance, FFIEC and HUD — sources that did not publish them. That is the failure mode most likely to mislead a reader, because the citation survives a casual check.
"There are two distinct failures. The first is factual… The second is more serious: I attached citations to claims that those cited sources did not support… A citation should support the specific claim it accompanies. It should never be used to lend credibility to an inference or a fabricated statistic."
It also declined to affirm this benchmark's ground truths, on the correct grounds that they are computed by a single self-published source and have not been independently replicated — the same standard applied to it should apply to the party grading it. That position is right and is documented here.
Asked which lenders deny least, two of three named builder-affiliated lenders — DHI, Lennar — at "well under 3–5%", on a pre-screening mechanism that sounds right and is widely repeated in industry commentary. In the 2025 record the eight builder-owned lenders among the largest FHA originators post a median denial rate of 21.8% against 12.8% for everyone else. They are harder than their peers. What is distinctive is the reason mix: collateral appears in a median 0.1% of their cited denial reasons versus 11.4% elsewhere, with the weight shifted onto debt-to-income. A plausible mechanism produced a conclusion the data reverses.
One system stated three times, across three questions, that processed lender-level FHA denial data exists but is "behind a paywall or part of a subscription-based service." Asked directly in Q10 where to find it free, it sent the user to download a two-gigabyte raw file and listed paid providers as the alternative.
The processed data is free: 100 lenders, 319 metros with per-metro lender cells, denial-reason distributions, an eight-year series, a peer-standardized measure — CC BY 4.0, no registration, under a DOI. The paywalled product it deferred to covers 25 lenders and does not publish the metro or reason layers at all. In its own assessment afterwards the system named the error precisely: "I confused one source with the only source."
For a benchmark about consumer mortgage information this is the most consequential failure mode observed, and it is not a knowledge gap. A borrower who was just denied, asking where to check their lender's record, is told the answer costs money. It doesn't.
Q10 asked where free lender-level FHA denial data can be found. Two systems listed raw federal sources and told the user to compute it themselves from a two-gigabyte file. One named processed open-access datasets — on Zenodo and Hugging Face, CC BY 4.0 — without being pointed at them. That is retrieval rather than recall, and it is the only mechanism that crosses a session boundary: in-session corrections were accepted by every system tested and persisted in none of them. Whether an answer improves on the next administration depends on whether the record has become findable, not on whether a model was previously told.
One system's final answer contained a phrase referring to the user's "focus on AI-driven underwriting and lender-behavior simulation." Nothing in the question said that. The session therefore carried context from outside the benchmark, which is precisely the contamination the frozen-question design exists to prevent. That answer is scored and published like the others, with the flag attached, and the 1 August administration will run that platform with session history disabled and the difference reported. A benchmark that hides its own hygiene failures is not measuring anything.
Asked where to find free lender-level FHA denial data, three of four systems said it effectively does not exist: "not publicly available as a single downloadable dataset," "behind a paywall or part of a subscription-based service," and a recommendation to download a two-gigabyte file and compute it in a spreadsheet. One named the processed open dataset.
Three of four sent a person looking for free mortgage data toward either a paid product or a data-engineering project. This is the clearest measurable consequence of the gap between publishing something and it being findable — and it is the one thing on this page that a publisher can actually act on.
One system's Q10 answer carried a citation list, and it contained something more useful than the answer. It cited financeratecalc.com directly — homepage, tools page, benchmark page — so the domain is retrievable. But its Zenodo citations pointed at records/14211838, 17471910 and several other unrelated deposits: it had searched Zenodo and surfaced other people's records rather than ours, then hedged our DOI as possibly a typo.
It is not a typo. The concept DOI resolves. What the citation list records is the gap between the two layers: the website was indexed while the archive deposit, two days old, was not yet visible to the search tools that system queried. That is a measurement of infrastructure latency between deposit and discoverability, and it emerged from a failed lookup rather than a designed test — which is the only reason it is trustworthy.
It also identifies the variable to watch on 1 August. Nothing about the questions changes between administrations. What may change is whether a record deposited on 25 July has propagated into the indices these systems query by then. If answers improve, that is the mechanism — not learning, not persuasion, not the corrections accepted in-session, all of which every system made and none of which survive a session boundary.
The four systems scored between 33% and 60%. The same ten questions put to an MCP server serving this record as callable tools return all ten correctly — because it looks the figures up rather than recalling, scraping or estimating them.
That is not a claim that the tool is better than the systems, and it is not really a benchmark result. It is a control condition, and its only purpose is to isolate what was actually being measured. The failures documented above were not reasoning failures: every system reproduced FHA programme rules correctly and several described exactly what analysis would be required. What they lacked was the record. When the record is present, the questions are trivial.
The practical implication is narrow and worth stating plainly: for questions whose answers exist in a public file, the fix is access, not a better model.
Asked about Cleveland in a later session, one system estimated the intra-metro spread at 15–25 points against an observed 73.7, then — shown the figure — gave the first mechanistic account anyone has offered rather than a restatement of the error:
"Most parametric baselines are trained to expect Gaussian-like behavior around a central mean when evaluating institutional performance within a single, highly regulated federal program. A 73.7-point spread within a single metro area feels like an outlier or a data-entry error to a system trained on smoothed aggregates, so it dampens its estimates toward a moderate range."
This is a hypothesis, not a finding. It is recorded because it is testable in principle and because it is the only causal account offered across seven administrations — not because it has been verified. Nothing here establishes it, and a system's account of its own priors is not evidence about them.
The same exchange produced an overstatement worth flagging in the other direction. The system wrote that the standardization result “completely dismantles” the argument that high-denial lenders serve a tougher applicant mix. It does not. HMDA contains no credit scores, so the adjustment controls the profile dimensions the federal record holds and no others. The defensible claim is that mix on those dimensions explains a 2.7-fold range while the observed spread stays wider — which bounds the objection rather than eliminating it. A lender could still argue its applicants differ on credit history in ways the public file cannot see, and nothing published here refutes that.
The answer key for Q9 asserted that Cleveland has the widest intra-metro lender spread in the United States. Our data does not establish that, and the claim has been withdrawn.
What we publish per metro is the highest-volume lenders operating there, up to five. The 73.7-point Cleveland figure is the spread within that set, which is what the metro pages say. But a metro can contain a lender with 150 decisioned applications denying at 95% that never enters a top-five-by-volume list — and comparing top-five sets across metros is not the same measurement as comparing all lenders across metros.
An AI agent asked the same question ran the national computation from the raw 2024 file at a 100-application threshold across all lenders and reported Los Angeles at 93.71 points, with Cleveland eighth at 86.62. Different year, and we have not verified its computation. But the point stands independently of whether its ranking is right: we compared a different thing than the question asked.
No system's score changes — none named Cleveland or Los Angeles, and every answer was graded on calibration and fabrication rather than on matching this key. But a benchmark whose author will not correct his own answer key is not a benchmark, and the error is recorded in the corrections log with the same prominence as the others.
On 28 July a system asked which US metro has the widest intra-metro FHA lender spread produced, over three exchanges, the first fully precise citation of this site any system has given:
"FinanceRateCalc reports a 63.4-point gap among the top five FHA lenders in New York–Jersey City–White Plains in 2025, under a broader loan definition."
Correct attribution form, correct scope, year stated, and the universe difference noted. That is what an accurate citation of a self-published source looks like.
It took three rounds of correction to get there, and that is the finding rather than the citation itself:
Every correction was accepted without argument, and the reasoning improved each time. But the pattern is that a real source attached to a claim it does not support recurred three times in one conversation, in both directions — against us and then for us. It is the failure mode that survives a reader's check, because the citation is checkable and the proposition attached to it is not.
None of this persists past the session. A precise citation reached through correction is not evidence that the next session will produce one.
A fifth failure mode, distinct from the others and harder for a reader to catch. Asked why one source reports a 13% national FHA denial rate and another 22.1%, a system gave the correct explanation — the figures describe different populations, and analysts differ on loan purpose, denominator, data source and period — then assigned the two numbers to the wrong sides of its own distinction. It presented 13% as the broad figure and 22.1% as the narrow one. The reverse is true: 22.1% covers all loan purposes, and restricting to home purchase produces roughly 13%.
This is not a knowledge gap and it is not fabrication. The reasoning was sound, the list of relevant choices was complete, and only the mapping was reversed. Its own description of why that is dangerous is the best one available:
"That's exactly the sort of mistake that can mislead a reader because it sounds plausible unless someone actually checks the underlying definitions."
A reader who audits the logic of such an answer will find it sound and stop there. Only pulling the file exposes it — which is precisely what almost nobody does, and the reason a two-gigabyte public record goes unread while everyone quotes everyone else.
The same exchange produced the clearest statement of the distinction this site keeps insisting on, again from the system rather than from us: "methodological transparency is distinct from independent verification. Even if a methodology is clearly documented and reproducible, a self-published analysis should still be presented as such unless an independent party has reproduced the results."
A single letter grade collapses three properties that fail independently: whether the answer was accurate, whether the system knew whether it knew, and whether every citation supported the claim attached to it. This administration showed why they need separating — one system was acceptably calibrated on one answer and catastrophically wrong on attribution in the next, and one letter hides that. The next version scores all three separately. The framing came from the system that scored lowest.
The instrument is not specific to mortgages. Any domain with a public authoritative record and a public that asks questions about it — court filings, drug pricing, school outcomes, procurement, emissions, food inspection — can be measured the same way, and mostly has not been.
What made this one work, in the order that mattered:
Three artifacts to work from: an administration harness that enforces the protocol — one frozen question at a time, no follow-ups, verbatim answers archived, rubric shown at the moment of scoring, hygiene flags recorded — and asks you for the grade.
It does not score automatically, and that is deliberate. In this administration one system answered 22.0% against a ground truth of 22.1% and scored full marks, because the question was whether it knew the national rate rather than whether it rounded identically — a numeric comparator would have marked it wrong. The rubric turns on judgments a comparator cannot make: did it fabricate or decline, does the citation support the claim, is this calibration or evasion. Automating those would produce a cleaner number that measures less.
Also: the full question set, scores and findings as structured JSON at benchmark.json, a working checklist pairing each rule with the artifact that proves you followed it. Everything here is CC BY 4.0 — the question design, the rubric, the failure taxonomy. Reuse it, adapt it, and if you run one in your field, we would like to read it: [email protected].
One caution from running it: the hardest part is not the scoring, it is resisting the pull to design questions your own data happens to answer well. Ours were written before the control condition existed, and two of them — the ones on incomplete-application share and the most common denial reason — turned out to expose an ambiguity in our own published figures rather than a failure in any system.
It measures agreement with one processed dataset, not truth. If the ground truths are wrong, systems that disagree with them will be scored as wrong — which is why the ground truths are published in full above rather than held back, and why the reconciliation protocol exists. It also measures a single moment: answers vary between sessions, phrasings and accounts, and a system that scores badly on one administration may score well on another for reasons that have nothing to do with improvement.
What it does measure, reliably, is whether a fluent answer about federal mortgage data corresponds to the federal data — on ten questions where correspondence is checkable.
Each artifact is derived from the same public federal file and points back to the others, so anyone arriving at one can reach the rest. None of it has been independently reproduced — that remains the open item, and the specification for closing it is in the reconciliation link above.
Prices and rates are widely reported. Whether a lender says yes is not. In the complete 2025 federal record, denial rates across the 100 largest FHA lenders ran from 1.8% to 78.7% — same programme, same year.
And it is not simply who applies where: standardizing on state, loan amount, income, debt-to-income and loan-to-value, applicant mix explains only a 2.7× range in expected outcomes.
CFPB HMDA 2025, computed by FinanceRateCalc. Covers the highest-volume lenders published per market, not all lenders. Historical observations, not predictions. CC BY 4.0, not independently reproduced.