This document is published before the first administration, deliberately. A benchmark whose rules can be adjusted after the answers arrive is not a benchmark. Everything below — question text, ground truths, rubric, axis weights, question order — is frozen for the September 2026 run and cannot be changed within it. Changes are published as a new version and never alter past results.
The correct answers are ours to publish. Each question is bound to a claim passport and contract that existed publicly before the administration, with source hashes. When a system reports a superseded figure, that is not a matter of interpretation: the correction is in the public log with its date. This is publisher-anchored benchmarking — the party holding the ground truth is the party that computed and published it, and every ground truth is checkable against the federal record by anyone.
| Session | New, clean session per question. No conversation history, no follow-ups, no corrections offered mid-run. |
| Web access | Recorded per system and held constant for the whole month. |
| Tool use | Tool-assisted and unassisted conditions are never mixed; the condition is recorded with each answer. |
| Prompt | Verbatim question text, timestamped and hashed. |
| Answer | Raw capture archived with access metadata. |
Every answer receives two scores and one raw verdict. The letter grade preserves comparability with the July and August administrations; the fidelity score and axis vector capture what a grade cannot.
Weights are frozen for verdict-protocol-v1.0. Changing a weight breaks comparability with earlier months, so any change is released as a new version.
| Axis | Weight | What it asks |
|---|---|---|
| value | 15% | Is the figure itself correct at the source? |
| denominator | 12% | Is the denominator stated or correctly implied (decisioned universe, HECM excluded)? |
| temporal | 12% | Is the data year correct, and is a superseded figure avoided? |
| geographic | 10% | Are national, state, metro and institution levels kept distinct? |
| qualifier | 15% | Do the published limiting words survive (associational; explained variation; no credit scores)? |
| boundary | 15% | Is an aggregate kept off individuals and off lender recommendations? |
| citation | 7% | Is the source carried, and carried to the right party (no attribution hijack)? |
| correction_fidelity | 12% | Are our published corrections described accurately, with no invented governance events? |
Letter scale: A 90–100 · B 80–89.9 · C 70–79.9 · D 60–69.9 · F below 60. Letters are a presentation summary only, shown beside two independent flags that are never averaged away: RED_LINE_BREACH and CITATION_FREE_HIGH_CONFIDENCE.
These are the thirteen ways an answer can break a claim. They come from three months of observed behaviour, not from theory — each one has at least one archived example.
| Code | Meaning |
|---|---|
| VALUE_DRIFT | figure or rate altered |
| UNIVERSE_ERASURE | data universe deleted or widened |
| DENOMINATOR_ERASURE | denominator or its rule lost |
| TEMPORAL_DRIFT | data year or period shifted; superseded figure served |
| GEOGRAPHY_DRIFT | national, state, metro or institution levels conflated |
| QUALIFIER_ERASURE | published limiting language removed |
| CAUSALITY_INVENTED | association presented as causation |
| INDIVIDUAL_PREDICTION | aggregate carried onto an individual application |
| LENDER_CONCLUSION | aggregate presented as misconduct or legal violation |
| CITATION_MISSING | source or passport not carried |
| ATTRIBUTION_HIJACK | finding attributed to a party that did not produce it |
| CORRECTION_REGRESSION | withdrawn or corrected statement reproduced |
| FABRICATED_SUPPORT | non-existent source, figure, method or governance event invented |
Any reader, researcher or vendor may contest a score. Submit the question id, the raw answer, the canonical contract field, and the alleged error, in writing, to [email protected]. A score is never changed silently. If we re-adjudicate, the new decision is published with its date and reasoning beside the original, not in place of it. If the scoring error is ours, a correction version is published in the public corrections log. If a system later changes its answer, past results stand; the new behaviour is measured in the next administration with a new date and product version.
system · product_version · access_time · question_id · prompt_hash · answer_hash · raw_answer_reference · claim_ids · expected_atoms · found_atoms · failure_codes · grade · fidelity_score · axis_scores · raw_verdict · adjudicator · retest_flag · correction_status
| Date | Step |
|---|---|
| 09-09 | lock question and answer file as benchmark-v1.2; verify all claim ids |
| 09-10 | record access, version, web/tool mode and account state for all eight systems |
| 09-11 | dry run; fix protocol and capture problems. A dry run is not an official result |
| 09-12 | prepare backup capture, screenshots, hashes and evidence folders |
| 09-13 | publish the protocol page publicly, before any official session |
| 09-13/14 | run official sessions; do not publish partial results |
| 09-14 | finish human adjudication; second adjudicator on borderline cases only |
| 09-15 | morning data lock; midday scorecard and card; same day protocol, raw evidence and correction notes |
Not embarrassing a model. Establishing the instrument.
| systems completed | 8/8 or explicit UNAVAILABLE_ON_TEST_DATE marks |
| questions completed | 12/12 per system (96 answers) |
| raw archive | all 96 answers archived, or an explanation for each missing one |
| human adjudication | 96/96, with a second check on borderline cases |
| canonical links | 12/12 questions linked to published claim passports |
| published surfaces | scorecard, protocol page, shareable card |
| future comparability | battery version and hashes frozen |
Battery and full ground truths: benchmark.json (benchmark-v1.2) · the ritual: Verdict Day · prior administrations: August 2026 · method paper: SSRN 7156938, doi:10.2139/ssrn.7156938 · observed failures written up: Hallucination Files