Item Audit

Your benchmark counts more questions than it measures with.

We ran item statistics over a published medical benchmark: 1 000 questions, 91 models. Across the whole field, 816 questions do some work. Across the top twenty models — the ones anybody is actually choosing between — 204 do.

The reliability score everyone quotes did not notice. It read 0.9944 on the full field and 0.9379 on the top twenty. Almost unchanged, while four fifths of the test stopped separating anyone.

experiments/PUBLIC-AUDIT-2026-10/RESULT.md

Choosing a model? The leaderboards cannot decide for you

Using Epoch AI's own published error bars, on 13 widely followed benchmarks, none of the nine neighbouring orderings in the top ten is statistically supported, and on 10 of them the first model cannot be separated from the fifth. A frontier leaderboard today is a leading group, not a ranking.

Verified model choice runs the leading models on your own tasks, compares them answer by answer — which a public leaderboard cannot do — and tells you which differences are real and which are noise. The evaluation it builds is itself checked: which of its questions separate the models, and which do not.

experiments/LEADERBOARD-REALITY-2026-10/RESULT.md — data: Epoch AI, CC BY 4.0; the analysis is ours

See a sample deliverable (in French): five coding models compared on 454 public tasks →

Fewer items, the same decisions

On SWE-bench Verified, 40% of the problems, chosen by looking only at the 20 earliest models, rank the 10 models released afterwards as the full benchmark does (Spearman 0.988), and all 20 supported pairwise orders keep their direction. A random 40% does clearly worse (median 0.878). Repeated on a benchmark of another kind (MedQA, 1 000 medical questions), half the items again rank the next 10 models as the full set does (0.985), though there a random half does almost as well. A selection ages as models improve, so it is refreshed.

experiments/COMPRESSION-SWEBENCH-2026-10/RESULT.md · experiments/COMPRESSION-MEDQA-2026-10/RESULT.md — both pre-registered

What an audit tells you

QuestionWhat the answer looks like
Which items are mis-keyed?Items your better models get wrong more often than your weaker ones. That is the signature of a wrong answer key.
Which items separate nobody?Every respondent answers them alike. You pay to run them and learn nothing.
How long is the test, really?Items counted against items carrying. On the benchmark above: 1 000 against 204.
Does the ranking hold at the top?Measurement error at the top of the leaderboard, against the gaps you are ranking on.

The problem is not ours to prove

A systematic review of 445 benchmarks in wide use, by 29 expert reviewers at the Oxford Internet Institute, found that 16.0% use an uncertainty estimate or a statistical test when comparing results. Three of its eight recommendations — use statistical methods to compare models, conduct an error analysis, justify construct validity — are what this audit does.

Measuring what Matters: Construct Validity in Large Language Model Benchmarks, NeurIPS 2025. Not our work, and not commissioned by us.

Price

 PriceTurnaround
Free scan — three numbers on your data0 €immediate
Verified model choice — the leading models run on your own tasks, and which differences are real4 900 €3 weeks
Benchmark claim check — does a published gap between two models hold?1 500 €48 hours
One benchmark, full audit1 900 €5 days
Your suite, up to 10 benchmarks7 500 €3 weeks
Shortened benchmark — the items that actually decide, so your evaluations run on a fraction of the items and reach the same conclusions; checked on models the selection never saw4 900 €3 weeks
Model watch — each time a new model is released, within 48 hours: is it really better on your tasks, and what would switching gain you? Item selection refreshed every quarter1 200 €/moongoing
Platform licence18 000 €/yrby contract

If a benchmark or suite audit finds neither a mis-keyed item nor redundancy above 20%, it is refunded. The clause does not apply to a model choice or a claim check: concluding that no gap is established is a result delivered, not a failure.

We do not know what is in your data. That uncertainty should sit with the seller, not with you.

What you should know before buying

Below 50 respondents, the screening list over-accuses. About one healthy item in seven gets flagged. We measured this on data whose defects we planted ourselves, and we publish the rate rather than waiting for you to find it. Above 50 respondents the strict list finds 88% to 100% of mis-keyed items with no false alarm in 500 replications.

experiments/DETECTION-2026-09/RESULT.md

Those detection rates come from synthetic data and are an upper bound. Real benchmarks break the independence assumptions behind them, so a real dataset can only be harder.

It needs per-item data, and you will probably have to extract it. Your harness already produces it — it cannot compute a score without knowing which items were right — but it may not be lying around in a file. We read four common harness formats directly to shorten that job.

scripts/from_harness.py

The code is public and the statistics are a century old

Item difficulty, discrimination, Cronbach's alpha, factor analysis, item response theory: all of it is public and has been for decades. Anyone can rebuild the scripts in a weekend.

What is sold is the reading — and the list of cases where the obvious behaviour is wrong, which was assembled the expensive way, by getting several of them wrong first and publishing the corrections.

Every number on this site, with the file it came from →

Start

The free check needs only a public leaderboard, and runs on your machine. Nothing is uploaded anywhere.

python scripts/leaderboard_check.py --table leaderboard.csv --items 1000

How to run it →  ·  What the paid audit adds →

To commission an audit, write to contact@itemaudit.fr.