Evidence
Every figure below names the file it comes from. Those files are in a public repository with the data and the commands that produced them.
How often the audit is right
Defects were planted in synthetic data — an item scored against the wrong answer, an item nearly everyone passes, an item unrelated to ability — and counted. The truth was known before the tool ran, so its accuracy is stated rather than assumed.
| Respondents | Mis-keyed found | Healthy falsely flagged | |
|---|---|---|---|
| 20 | screening | 0.976 | 0.144 |
| asserting | 0.500 | 0.002 | |
| 50 | screening | 0.999 | 0.051 |
| asserting | 0.880 | 0.000 | |
| 200 | screening | 1.000 | 0.003 |
| asserting | 1.000 | 0.000 |
Detection is never quoted without false alarms. A tool that flags every item detects every defect; one that flags nothing never raises a false alarm. Either number alone recommends the opposite choice.
The limit of this result. Responses came from a two-parameter logistic model with independent draws. Real benchmarks break that: items share topics and respondents share training data. So these are an upper bound on detection and a lower bound on false alarms. That sentence travels inside the data file as well as here, because a bound separated from its condition gets quoted as a certainty.
Checked against human experts
The flags find real key errors. On 5 005 MMLU items relabelled by experts (MMLU-Redux 2.0, CC BY 4.0), the 7% of items on the screening list hold 54% of the expert-found key errors: 11.1% of flagged items carry one, against 0.8% of the rest (×14.6). On the strict list the rate is 21.9%. All four pre-registered predictions held. Most flagged items are still fine: a flag says what to read first.
experiments/VALIDATION-REDUX-2026-10/RESULT.mdA cheap model can say what is probably wrong. Blind to the flag, deepseek-v4.1-flash judged the 371 flagged items. Among those it called mis-keyed, 43% carry an expert-found key error, against 11% for the flag alone; it named 24 of the 41. Repeated on 300 questions written the same day, with a wrong key planted in 60, it caught 59, named the right answer each time, and called none of the 240 clean questions wrong, so the first result does not rest on memorised public data. Planted errors are easier than natural ones: quote the 43%.
experiments/DIAGNOSIS-REDUX-2026-10/RESULT.md · experiments/DIAGNOSIS-UNSEEN-2026-10/RESULT.mdLeaderboards, read with their own error bars
On 13 benchmarks published by Epoch AI with a standard error per model, none of the nine neighbouring orderings in the top ten is supported, and on 10 of them the first model cannot be separated from the fifth. On release day, GPT-6.1 Sol was first alone on 1 of 9 comparable benchmarks, inside the error bars.
experiments/LEADERBOARD-REALITY-2026-10/RESULT.md · experiments/LAUNCH-CHECK-2026-10/RESULT.mdTwo public benchmarks, audited
A medical benchmark: 1 000 items, 91 models
| Items carrying | Alpha | Running backwards: strict / screening | |
|---|---|---|---|
| All 91 models | 816 / 1000 | 0.9944 | 25 / 76 |
| Top 20 models | 204 / 1000 | 0.9379 | 7 / 138 |
These items are not claimed to be mis-keyed. An item that the stronger models get wrong is usually a wrong key — on MMLU, all five checked by hand were. On this medical benchmark they have not been checked, the audit itself says they do not look mis-keyed, and settling it needs clinical competence we do not have. At twenty respondents the screening count also includes about one healthy item in seven, which is why the strict count is so much smaller.
The standard reliability score missed this. It moved from 0.9944 to 0.9379 while four fifths of the test stopped separating anyone. The mean correlation between items fell by a factor of ten over the same span; a thousand items manufacture a high alpha out of almost nothing. On the top-20 panel, 875 of the 1 000 items raise alpha when dropped.
This audit was pre-registered. Seven predictions were written and committed before the data was fetched; two failed. The failures are published unchanged, and one of them corrected an earlier claim of ours. experiments/PUBLIC-AUDIT-2026-10/PREREGISTRATION.md
A knowledge benchmark: 111 items, 91 models
| Items carrying | Alpha | Running backwards: strict / screening | |
|---|---|---|---|
| All 91 models | 72 / 111 | 0.948 | 6 / 9 |
| Top 20 models | 4 / 111 | −0.442 | 2 / 18 |
A negative alpha is not a small reliability. It is negative average covariance between items: they contradict each other more than chance would produce, so no coherent scale is left on which to put one respondent ahead of another.
Precision at the top of the leaderboard
| Where on the leaderboard | Information | Measurement error |
|---|---|---|
| Median model | 24.87 | 0.20 |
| Median of the top 20 | 5.99 | 0.41 |
| Strongest model | 0.17 | 2.40 |
The models being chosen between are the ones the test separates least.
What the tool refuses to tell you
Every vendor shows what their tool outputs. The other list is longer and more useful.
- A graded score is refused, not rounded. Rounding a 1-to-5 rubric once reported a healthy instrument as measuring with nothing — silently.
- Alpha is undefined rather than zero when nothing varies. Zero says the test is unreliable; undefined says your data cannot tell you.
- A trial missing an item is dropped, not scored zero. Filling a zero turns an absent answer into a wrong one.
- An unmeasured field reports “unmeasured”, never 0 and never ok.
- There is no overall score, deliberately. An average lets a strong field hide a missing one.
- Item response theory answers “the question cannot be answered at this sample size” when the test has no power to distinguish the hypotheses.
- A factor resting on one respondent is flagged as resting on one respondent.
- A threshold is put against the interval, not the estimate — a refusal to assert from twenty respondents what twenty respondents cannot support.
Six times this work proved its own authors wrong
The fear anyone brings to an audit is that the auditor will say what the buyer wants to hear. There is no argument against that fear — only a record.
- The tool contradicted its author. We recommended deleting four contested questions. Item analysis showed they were the four highest-discriminating items in the instrument. Deleting them would have removed most of what the test measured while leaving a page of numbers that still looked healthy. Pinned in a test so it cannot quietly disappear.
- A quiz generator keyed three questions to the opposite of its own document. No correct option existed. The reader answered “the text does not say” in 24 runs of 24 and was marked wrong every time. It was right every time.
- A benchmark paired its arms by list position, so one failure silently turned a paired confidence interval into an unpaired one — still printed, still narrow, no longer meaning what it said.
- A healthy instrument was declared dead, silently, by a rounding rule. Alpha 0.864 with 7 of 8 items carrying was reported as “a test of 8 items that measures with 0”.
- An accusation against our own data was published, then retracted. So was a second claim, contradicted within the hour by a free check on data we already held.
- The accuracy work found flaws in the accuracy claims it was built to support — the 14% false alarm rate above, and a detection rate that does not improve monotonically with sample size.
Four of those six are defects we shipped. An auditor with no such record is not necessarily worse. There is simply no way to tell.
Not our work, and not commissioned by us
A systematic review of 445 benchmarks in wide use, 29 expert reviewers, Oxford Internet Institute, NeurIPS 2025:
| Of 445 benchmarks | Share |
|---|---|
| Use an uncertainty estimate or statistical test to compare results | 16.0% |
| Present evidence for construct validity | 53.4% |
| Define the phenomenon they measure | 78.2% |
| …of those, a definition that is contested rather than agreed | 47.8% |