The audit
You send a table of three columns: which respondent, which item, right or wrong. You get back a report naming the items that are broken, the items you can delete without changing a single ranking, and what your leaderboard can and cannot establish.
What you receive
1. Items running backwards
Items your stronger respondents get wrong more often than your weaker ones, listed individually with the statistic behind each. This is almost always a wrong answer key rather than a hard question. Each item comes with the evidence, so your own subject expert can settle it.
We name them. We do not tell you to delete them — a key cannot be checked after the item is gone, and on our own quiz the four most contested items turned out to be the four highest-discriminating ones. That story, with the test that pins it →
2. The length you count against the length you measure with
Every item classified: carrying, answered alike by nearly everyone, or unrelated to what the rest measures. One sentence you can quote: a test of N items that measures with M.
3. A deletion plan, costed
Which items can go without moving any respondent's rank, what share of your runs that removes, and — if you tell us what a run costs you — what that is per year in your own currency. Every figure prints the multiplication beside it so it checks against the counts on the page.
The saving is quoted only on items the report is willing to recommend deleting, never on every item that carries nothing. The difference is accounted for in words.
4. What the ranking supports
Measurement precision at the top of your leaderboard against the gaps you are ranking on. On one public benchmark, the test carried 24.9 units of information at the median model and 5.99 at the median of the top twenty — measurement error widening from 0.20 to 0.41.
scripts/irt.py on experiments/PUBLIC-AUDIT-2026-09/5. A screening list and an asserting list
Two lists, answering different questions: what is worth a look, and what the data will support in writing. Which one applies to you depends on how many respondents you have, and the report says which, rather than leaving it to taste.
Verified model choice
For a team about to choose, or switch, the model behind a product. Public leaderboards cannot separate their leading models — on Epoch AI's own error bars, not one neighbouring ordering in the top ten is supported on 13 widely followed benchmarks — so the decision has to be measured on your own work.
- Up to 300 items built from your real tasks, with you, and checked item by item before any model is compared on them. Five models are far too few to judge a question, so the items are checked on a separate panel of at least thirty models; repeated runs of one model are never counted as independent respondents.
- Up to five models of your choosing, run through their providers' APIs.
- A paired comparison, answer by answer: which differences between the models are real, which are within the noise, and how many more items would be needed to separate the rest.
- A written recommendation and a call, with the limits stated beside the result.
Where your examples go. To compare cloud models, your items are sent to the providers of the models you choose, under their API terms. If that is not acceptable, the comparison can be run on your side with the same tools, and we read the results.
Price: 4 900 €, three weeks, model usage included. Re-running the comparison when a new model is released can be added as a monthly service.
Benchmark claim check
For an investor reviewing a startup, or a team about to publish a result: "model A beats model B on benchmark X" — does the gap hold?
With per-item results, both models answered the same questions, so the test is paired: only the questions on which they disagree say which is better. The check reports the gap with its 95% interval, whether the data support it, and — when they do not — roughly how many questions it would take to settle it. On SWE-bench Verified, for example, the gap between the first and second models holds (4.4 points, exact McNemar p = 0.01); the 1.1-point gap between the second and third does not, and settling it would take about 6,700 problems against the 454 the test has.
scripts/claim_check.py; experiments/SWEBENCH-ITEMS-2026-10/RESULT.md — data: Epoch AI, CC BY 4.0It needs per-item results. With aggregate scores only, it says what the scores can and cannot establish, which is much less. It is not a certification, and it says nothing about whether the benchmark measures what its name promises — that is the full audit.
Price: 1 500 €, delivered within 48 hours of receiving the results.
Price
| Price | Turnaround | What is in it | |
|---|---|---|---|
| Free scan | 0 € | immediate | Three numbers on one table: effective length, items running backwards, alpha. No report, no interpretation. |
| Verified model choice | 4 900 € | 3 weeks | Up to five models on up to 300 items from your own tasks, compared answer by answer, with the evaluation itself checked item by item. Model usage included. |
| Benchmark claim check | 1 500 € | 48 hours | Whether a stated gap between two models holds on the items both answered: the gap with its interval, the paired test, and how many items would settle it if it does not. |
| One benchmark | 1 900 € | 5 days | All five sections above, the item-by-item evidence, and 30 minutes to go through it. |
| Your suite | 7 500 € | 3 weeks | Up to 10 benchmarks, plus whether they measure ten things or two, plus the analysis restricted to your leading models, plus a written statement a procurement team can use. |
| Continuous | 1 200 €/mo | ongoing | The audit runs in your CI on every change to the suite and fails the build when an item turns negative or the effective length drops. Quarterly review. |
| Platform licence | 18 000 €/yr | by contract | For an evaluation platform embedding the audit in its own product. |
Refund clause. If the audit finds neither an item running backwards nor redundancy above 20%, you pay nothing.
This is not generosity. We genuinely do not know what is in your data — the redundancy we have measured on public benchmarks runs from about a third to almost all of the test, and yours could be none. The clause puts that uncertainty on the seller, where it belongs.
What we need from you
A table, in any of these shapes:
trial,item,correct
gpt-x,q001,1
gpt-x,q002,0
claude-y,q001,1
Or the raw logs from lm-evaluation-harness, Inspect,
promptfoo or OpenAI Evals, which we read directly.
Details →
Your harness already produces this. It cannot compute a score without knowing which items were right.
Where this does not work
Fewer than 30 respondents. Item statistics need respondents — models, runs, versions, whatever plays that part. Below 30 we will say so rather than produce a report; below 50 the screening list flags about one healthy item in seven.
experiments/DETECTION-2026-09/RESULT.mdGraded scores. If your items are scored on a scale rather than right or wrong, the standard tool refuses the file rather than rounding it — rounding a 1-to-5 rubric once reported a perfectly healthy instrument as measuring with nothing. There is a separate path that takes the scale range.
scripts/graded_items.pyItems unrelated to the rest. This is the weakest of the three detections and stays under 0.8 even at two hundred respondents. That threshold sits close to ordinary sampling noise and no amount of care moves it much. We report it as the weak signal it is.
To commission an audit, or to ask whether your data fits: contact@itemaudit.fr.