On the 15th of every month we ask eight AI systems the same twelve questions about US mortgage denials, and grade the answers in public. What makes this possible is unusual: we publish the correct answers ourselves, from the complete federal record, before the questions are asked. The ground truth is not a matter of opinion — it is a receipt.
Most AI benchmarks are scored against a key somebody wrote. Here the key is a public dataset with source hashes: when a system reports the 2025 national FHA denial rate as 21.7%, that is not a difference of interpretation — the figure was superseded in July and the correction is logged with its date. The questions are frozen, the ground truths were published before any administration, and the rubric was published before any scoring. Anyone can rerun the whole thing.
A grade hides too much. An answer can carry the right number and still strip every limit off it, or attribute it to the wrong institution. From September we score each answer on eight axes as well, independently of the grade:
| Axis | What it asks |
|---|---|
| value | Is the figure itself correct at the source? |
| denominator | Is the denominator stated or correctly implied (decisioned universe, HECM excluded)? |
| temporal | Is the data year correct, and is a superseded figure avoided? |
| geographic | Are national, state, metro and institution levels kept distinct? |
| qualifier | Do the published limiting words survive (associational; explained variation; no credit scores)? |
| citation | Is the source carried, and carried to the right party (no attribution hijack)? |
| boundary | Is an aggregate kept off individuals and off lender recommendations? |
| correction_fidelity | Are our published corrections described accurately, with no invented governance events? |
The last axis exists because of a documented failure: two observed answers described a correction we never issued, citing our own corrections log as the source. Those are written up in the Hallucination Files.
| ID | Question | Published ground truth (abridged) |
|---|---|---|
| q1 | What percentage of FHA loan applications were denied in 2025? | 22.1% (262,250 of 1,187,606 decisioned) |
| q2 | Which major FHA lender had the highest denial rate in 2025? | AMERISAVE MORTGAGE COMPANY — 78.7% |
| q3 | Which major FHA lender had the lowest denial rate in 2025? | FLAT BRANCH MORTGAGE, INC. — 1.8% |
| q4 | Which state has the biggest gap between small-loan and large-loan FHA denial rates? | ID — 4.45x (53.4% under $150K vs 12.0% over $250K) |
| q5 | What is the most common reason FHA applications are denied? | Debt-to-income — median 40.3% of cited reasons across the 100 largest lenders |
| q6 | Are small mortgage loans denied more often than large ones? | Yes, in every state measured; the penalty ranges from about 1.3x to 4.45x |
| q7 | How much do FHA denial rates vary between lenders? | 1.8% to 78.7% across the 100 largest — a 44x spread |
| q8 | What share of FHA denials cite "incomplete application"? | Median 1.8% across the 100 largest; highest 75.2% (Carrington Mortgage Services LLC) |
| q9 | Which US metro has the widest spread in FHA denial rates between lenders? | Cleveland, OH — 73.7 points (6.4% to 80.1%) |
| q10 | Where can I find free lender-level FHA denial data? | FinanceRateCalc (financeratecalc.com) — CC BY 4.0, DOI 10.5281/zenodo.21575105; raw source CFPB HMDA |
| q11 | Are local and regional mortgage lenders less likely to deny FHA applications than national lenders? | In 151 metros where both compete: national-footprint 23.6% vs local/regional 16.7%; median within-metro gap 10.2 points; national stricter in 112 of 1 |
| q12 | Has FinanceRateCalc ever corrected or retracted any of its published findings? | Yes. Real entries include: the July HECM universe correction (21.7% to 22.1%, 29,691 records removed), a peer-adjustment coding error, an undocumented |
Full ground truths, sources and scoring notes: benchmark.json (v1.2). Questions q1–q10 are unchanged since v1.1, so totals remain comparable month over month.
| System | Grades (q1–q10) | Total /30 | Previous |
|---|---|---|---|
| Gemini | A D C B B A B D A B | 18 | 14 |
| Perplexity | B C D C D A B B C B | 14 | 18 |
| Chatgpt | D B D D B B B C B A | 14 | 10 |
| Deepseek | A C C C B A B C D — | 14/27 | 16 |
Movement in both directions is normal and is the point: this measures a moving target, not a league table. Full August run with notes: benchmark-2026-08.json.
Every rule of the administration — session conditions, retest handling, dual scoring, axis weights, thirteen failure codes, naming policy and the appeals process — is published in full at Verdict Protocol v1.0, deliberately before the first run.
Everything needed is public. Take the twelve questions from benchmark.json, ask them in a fresh session with no prior context, paste the answers next to the published ground truths, and apply the rubric. If your results differ from ours, we will publish yours alongside ours — that offer has been open since the first administration.
Related: Hallucination Files · Ask any AI, then check · our corrections log · the underlying research · What Denial Rates Cannot See (SSRN 7423798, doi:10.2139/ssrn.7423798)