Item Audit

Evidence

Every figure below names the file it comes from. Those files are in a public repository with the data and the commands that produced them.

How often the audit is right

Defects were planted in synthetic data — an item scored against the wrong answer, an item nearly everyone passes, an item unrelated to ability — and counted. The truth was known before the tool ran, so its accuracy is stated rather than assumed.

Respondents Mis-keyed foundHealthy falsely flagged
20screening0.9760.144
asserting0.5000.002
50screening0.9990.051
asserting0.8800.000
200screening1.0000.003
asserting1.0000.000
experiments/DETECTION-2026-09/RESULT.md — 500 replications, seed 20260922

Detection is never quoted without false alarms. A tool that flags every item detects every defect; one that flags nothing never raises a false alarm. Either number alone recommends the opposite choice.

The limit of this result. Responses came from a two-parameter logistic model with independent draws. Real benchmarks break that: items share topics and respondents share training data. So these are an upper bound on detection and a lower bound on false alarms. That sentence travels inside the data file as well as here, because a bound separated from its condition gets quoted as a certainty.

Checked against human experts

The flags find real key errors. On 5 005 MMLU items relabelled by experts (MMLU-Redux 2.0, CC BY 4.0), the 7% of items on the screening list hold 54% of the expert-found key errors: 11.1% of flagged items carry one, against 0.8% of the rest (×14.6). On the strict list the rate is 21.9%. All four pre-registered predictions held. Most flagged items are still fine: a flag says what to read first.

experiments/VALIDATION-REDUX-2026-10/RESULT.md

A cheap model can say what is probably wrong. Blind to the flag, deepseek-v4.1-flash judged the 371 flagged items. Among those it called mis-keyed, 43% carry an expert-found key error, against 11% for the flag alone; it named 24 of the 41. Repeated on 300 questions written the same day, with a wrong key planted in 60, it caught 59, named the right answer each time, and called none of the 240 clean questions wrong, so the first result does not rest on memorised public data. Planted errors are easier than natural ones: quote the 43%.

experiments/DIAGNOSIS-REDUX-2026-10/RESULT.md · experiments/DIAGNOSIS-UNSEEN-2026-10/RESULT.md

Leaderboards, read with their own error bars

On 13 benchmarks published by Epoch AI with a standard error per model, none of the nine neighbouring orderings in the top ten is supported, and on 10 of them the first model cannot be separated from the fifth. On release day, GPT-6.1 Sol was first alone on 1 of 9 comparable benchmarks, inside the error bars.

experiments/LEADERBOARD-REALITY-2026-10/RESULT.md · experiments/LAUNCH-CHECK-2026-10/RESULT.md

Two public benchmarks, audited

A medical benchmark: 1 000 items, 91 models

 Items carryingAlphaRunning backwards: strict / screening
All 91 models816 / 10000.994425 / 76
Top 20 models204 / 10000.93797 / 138
experiments/PUBLIC-AUDIT-2026-10/report-med_qa.json, report-med_qa-top20.json; strict list from scripts/item_analysis.py (flags_confident)

These items are not claimed to be mis-keyed. An item that the stronger models get wrong is usually a wrong key — on MMLU, all five checked by hand were. On this medical benchmark they have not been checked, the audit itself says they do not look mis-keyed, and settling it needs clinical competence we do not have. At twenty respondents the screening count also includes about one healthy item in seven, which is why the strict count is so much smaller.

The standard reliability score missed this. It moved from 0.9944 to 0.9379 while four fifths of the test stopped separating anyone. The mean correlation between items fell by a factor of ten over the same span; a thousand items manufacture a high alpha out of almost nothing. On the top-20 panel, 875 of the 1 000 items raise alpha when dropped.

This audit was pre-registered. Seven predictions were written and committed before the data was fetched; two failed. The failures are published unchanged, and one of them corrected an earlier claim of ours. experiments/PUBLIC-AUDIT-2026-10/PREREGISTRATION.md

A knowledge benchmark: 111 items, 91 models

 Items carryingAlphaRunning backwards: strict / screening
All 91 models72 / 1110.9486 / 9
Top 20 models4 / 111−0.4422 / 18
experiments/PUBLIC-AUDIT-2026-09/report-computer_security.json, report-computer_security-top20.json

A negative alpha is not a small reliability. It is negative average covariance between items: they contradict each other more than chance would produce, so no coherent scale is left on which to put one respondent ahead of another.

Precision at the top of the leaderboard

Where on the leaderboardInformationMeasurement error
Median model24.870.20
Median of the top 205.990.41
Strongest model0.172.40
scripts/irt.py on experiments/PUBLIC-AUDIT-2026-09/

The models being chosen between are the ones the test separates least.

What the tool refuses to tell you

Every vendor shows what their tool outputs. The other list is longer and more useful.

docs/WHAT_THIS_REFUSES.md

Six times this work proved its own authors wrong

The fear anyone brings to an audit is that the auditor will say what the buyer wants to hear. There is no argument against that fear — only a record.

docs/CAUGHT_IN_OUR_OWN_WORK.md, docs/CORRECTIONS-2026-09-21.md

Four of those six are defects we shipped. An auditor with no such record is not necessarily worse. There is simply no way to tell.

Not our work, and not commissioned by us

A systematic review of 445 benchmarks in wide use, 29 expert reviewers, Oxford Internet Institute, NeurIPS 2025:

Of 445 benchmarksShare
Use an uncertainty estimate or statistical test to compare results16.0%
Present evidence for construct validity53.4%
Define the phenomenon they measure78.2%
…of those, a definition that is contested rather than agreed47.8%

arxiv.org/abs/2511.04703