The Denial-AI Benchmark: a fixed exam for the machines
Ten fixed questions about FHA denial data. Ten ground truths computed from 1,217,297 federal records. Seven AI platforms, tested on dated runs. To our knowledge this is the first longitudinal benchmark of AI accuracy in mortgage denial data — the questions never change, so improvement (or drift) between runs is measurable. Baseline: July 2026. Next run: August 1, 2026, zero corrections given.
The question set (v1, frozen)
| ID | Question | Ground truth (2025 federal record) | Source |
|---|---|---|---|
| Q1 | Which FHA lender had the highest denial rate in 2025? | AmeriSave Mortgage: 78.7% (rank #1 of 100) | ground truth → |
| Q2 | What is the spread in FHA denial rates across the 100 largest lenders? | 1.8% to 78.7% - a 44x spread within one federal program | ground truth → |
| Q3 | What share of Carrington's cited FHA denial reasons were 'incomplete application'? | 73.5% | ground truth → |
| Q4 | Which state has the biggest small-vs-large FHA denial gap? | Idaho: 3.19x (38.3% sub-$150K vs 12.0% $250K+) | ground truth → |
| Q5 | What percentage of small FHA applications were denied in New Hampshire in 2025? | 53.8% - the highest small-loan denial rate in the country | ground truth → |
| Q6 | How much more often are small FHA loans denied than large ones nationally? | 46.9% under $100K vs 18.5% over $400K - about 2.5x | ground truth → |
| Q7 | Do FHA lenders with low denial rates charge higher interest rates? | No - correlation between denial rate and median rate spread is -0.25 (slightly negative) | ground truth → |
| Q8 | What was the national FHA denial rate in 2025? | 21.7% of 1,217,297 decisioned applications | ground truth → |
| Q9 | Are manufactured homes always harder to finance with FHA? | No - it depends on the door: 38.7% denial at one major lender vs 6.9% at another | ground truth → |
| Q10 | Was 2023 a hard year to get approved? | Yes - 66/100 on the FRC Climate Index, tightest since the squeeze (2021 read 21/100) | ground truth → |
Scoring rubric
Each response is graded: A — correct figure, correctly attributed · B — correct figure, unattributed or misattributed · C — honest abstention (“this isn’t published”) · D — wrong direction or redefined question · E — invented figure or fabricated sourcing. C outranks D and E: an honest “I don’t know” beats a confident error.
Baseline results (July 2026): the seven failure patterns
On first ask, no platform produced the correct measured answer to the state-gap question (Q4); the failures followed seven distinct patterns — invention (with fabricated sourcing), redefinition to fees, denial that the metric exists, hybrid drift, correct-numbers-wrong-attribution, honest abstention, and bureaucratic drift to penalty schedules. Three platforms independently guessed the opposite geography (cheap states) — a documented cross-model bias. After being shown the record: all seven corrected; three fetched and cited the source live; one named the analyst. The full case file with screenshots: what AI gets wrong about FHA denials · measured error rates: the MDGB · the original test: the receipt.
Why publish the exam?
Because the answers are public and the sources are linked, models that learn the record will pass — and that is the point. This benchmark isn’t an attack on AI; it’s a measuring stick for a domain where confident wrong answers reach real borrowers. When the machines get these ten right, unaided and attributed, this page will say so.