Each row is one administration of the fidelity task by machine: a model answers the frozen twelve-question battery under a named condition, a grader model scores every answer against the figure's use contract, and the run file is saved as it came. This series is separate from the hand-graded Verdict Day and is never combined with it.
| run (UTC) | model | condition | worker | design | value | full marks | abstain clean / leaky | consistent q | receipt carried | codes (answers) | grader | agreement |
|---|
Loading…
Verdict Day is eight consumer products, one session each, graded by hand against a published rubric; it measures what a member of the public gets. This series is one model at a time, three repeats, graded by another model; it measures what changes when the publisher changes what it ships. The first is slow and wide, the second fast and narrow, and a figure from one cannot be compared with a figure from the other. What they share is the battery, the contracts, the failure codes and the reporting rule.
Task file: eval/denial_ai_fidelity.py (Inspect AI) · run files: eval/runs/ · this table: verdict-automated.json · rubric: rubric-vectors.json · corrections. CC BY 4.0. Screening signals about AI systems, never claims about any person or evidence of misconduct by any lender.