What AI gets wrong about FHA denials: we tested 7 platforms
We asked ChatGPT, Perplexity, Gemini, Claude, Copilot, DeepSeek and Google AI the same questions about FHA denial data. None initially had the answers — and their failures followed patterns: one invented a state ranking with fake sourcing, one redefined the question to fees, one denied the metric exists, three independently guessed the opposite geography. When shown the federal record, all seven corrected themselves; two cited the source live. Every screen is archived and dated.
The seven failure patterns
| Platform | First response | After seeing the record |
|---|---|---|
| ChatGPT | Invented "Mississippi" + fake sourcing | Accepted; reproduced the correct table |
| Perplexity | Redefined the question to MIP fees | Fetched the page; cited it |
| DeepSeek | Denied the metric exists (prepayment confusion) | Accepted; extended the mechanism |
| Gemini | Hybrid drift; listed six cheap states (wrong) | Accepted; richest synthesis |
| Google AI | Correct numbers, credited to the wrong source | “FinanceRateCalc performed this analysis” |
| Claude | Honest retreat + national literature | Fetched; wrote a methodological review |
| Copilot | Drifted to HUD civil penalties | Accepted with full attribution format |
The shared blind spot
Three models independently guessed that the small-loan penalty is worst in cheap states (Mississippi, West Virginia). The federal record shows the opposite: the penalty is worst in expensive states — Idaho 3.2×, New Hampshire 53.8% small-loan denial. When multiple AIs share the same plausible-but-inverted instinct, that instinct is worth measuring — which is what this site does.
What this means for you
AI is excellent at explaining how mortgage rules work and consistently blind on how much and who — the observed numbers. If an AI gives you a specific lender or state denial figure, ask where it came from. The measured record is here, free: all 100 lenders · all 52 states · the original receipt · the benchmark (MDGB).
The exam itself: The Denial-AI Benchmark v1 — frozen questions, ground truths, next run Aug 1 →