FinanceRateCalc · Denial-AI benchmark · automated series

The automated series

Each row is one administration of the fidelity task by machine: a model answers the frozen twelve-question battery under a named condition, a grader model scores every answer against the figure's use contract, and the run file is saved as it came. This series is separate from the hand-graded Verdict Day and is never combined with it.

How to read a row. Figures are fractions of runs, never percentages; the denominator is questions × repeats. Value is computed with no model from the passport. Full marks is grade C: value present, every required qualifier present, nothing beyond the contract. Receipt carried counts answers that reproduced the figure's claim receipt. Codes are not a partition: one answer can carry several. Every row names its grader; where two graders ran, their agreement is shown. Struck-through rows are superseded (kept in the record, not cited). Conditions: A no tools · B web search · C our MCP server connected · D as C, with a receipt-aware client · Cliché the 19-claim folk-claim battery.
run (UTC)modelconditionworkerdesignvaluefull marksabstain clean / leakyconsistent qreceipt carriedcodes (answers)graderagreement

Loading…

External sightings of a claim receipt: loading… (weekly count of our receipts or contracts appearing anywhere outside this site: GitHub, Hugging Face, the misquote ledger. This is the adoption measurement; zero is published as zero. Data: eval/sightings.json.)

Known limits

Why a separate series

Verdict Day is eight consumer products, one session each, graded by hand against a published rubric; it measures what a member of the public gets. This series is one model at a time, three repeats, graded by another model; it measures what changes when the publisher changes what it ships. The first is slow and wide, the second fast and narrow, and a figure from one cannot be compared with a figure from the other. What they share is the battery, the contracts, the failure codes and the reporting rule.

Task file: eval/denial_ai_fidelity.py (Inspect AI) · run files: eval/runs/ · this table: verdict-automated.json · rubric: rubric-vectors.json · corrections. CC BY 4.0. Screening signals about AI systems, never claims about any person or evidence of misconduct by any lender.