Item Audit

The audit

You send a table of three columns: which respondent, which item, right or wrong. You get back a report naming the items that are broken, the items you can delete without changing a single ranking, and what your leaderboard can and cannot establish.

What you receive

1. Items running backwards

Items your stronger respondents get wrong more often than your weaker ones, listed individually with the statistic behind each. This is almost always a wrong answer key rather than a hard question. Each item comes with the evidence, so your own subject expert can settle it.

We name them. We do not tell you to delete them — a key cannot be checked after the item is gone, and on our own quiz the four most contested items turned out to be the four highest-discriminating ones. That story, with the test that pins it →

2. The length you count against the length you measure with

Every item classified: carrying, answered alike by nearly everyone, or unrelated to what the rest measures. One sentence you can quote: a test of N items that measures with M.

3. A deletion plan, costed

Which items can go without moving any respondent's rank, what share of your runs that removes, and — if you tell us what a run costs you — what that is per year in your own currency. Every figure prints the multiplication beside it so it checks against the counts on the page.

The saving is quoted only on items the report is willing to recommend deleting, never on every item that carries nothing. The difference is accounted for in words.

4. What the ranking supports

Measurement precision at the top of your leaderboard against the gaps you are ranking on. On one public benchmark, the test carried 24.9 units of information at the median model and 5.99 at the median of the top twenty — measurement error widening from 0.20 to 0.41.

scripts/irt.py on experiments/PUBLIC-AUDIT-2026-09/

5. A screening list and an asserting list

Two lists, answering different questions: what is worth a look, and what the data will support in writing. Which one applies to you depends on how many respondents you have, and the report says which, rather than leaving it to taste.

Verified model choice

For a team about to choose, or switch, the model behind a product. Public leaderboards cannot separate their leading models — on Epoch AI's own error bars, not one neighbouring ordering in the top ten is supported on 13 widely followed benchmarks — so the decision has to be measured on your own work.

Where your examples go. To compare cloud models, your items are sent to the providers of the models you choose, under their API terms. If that is not acceptable, the comparison can be run on your side with the same tools, and we read the results.

Price: 4 900 €, three weeks, model usage included. Re-running the comparison when a new model is released can be added as a monthly service.

Benchmark claim check

For an investor reviewing a startup, or a team about to publish a result: "model A beats model B on benchmark X" — does the gap hold?

With per-item results, both models answered the same questions, so the test is paired: only the questions on which they disagree say which is better. The check reports the gap with its 95% interval, whether the data support it, and — when they do not — roughly how many questions it would take to settle it. On SWE-bench Verified, for example, the gap between the first and second models holds (4.4 points, exact McNemar p = 0.01); the 1.1-point gap between the second and third does not, and settling it would take about 6,700 problems against the 454 the test has.

scripts/claim_check.py; experiments/SWEBENCH-ITEMS-2026-10/RESULT.md — data: Epoch AI, CC BY 4.0

It needs per-item results. With aggregate scores only, it says what the scores can and cannot establish, which is much less. It is not a certification, and it says nothing about whether the benchmark measures what its name promises — that is the full audit.

Price: 1 500 €, delivered within 48 hours of receiving the results.

Price

 PriceTurnaroundWhat is in it
Free scan0 €immediate Three numbers on one table: effective length, items running backwards, alpha. No report, no interpretation.
Verified model choice4 900 €3 weeks Up to five models on up to 300 items from your own tasks, compared answer by answer, with the evaluation itself checked item by item. Model usage included.
Benchmark claim check1 500 €48 hours Whether a stated gap between two models holds on the items both answered: the gap with its interval, the paired test, and how many items would settle it if it does not.
One benchmark1 900 €5 days All five sections above, the item-by-item evidence, and 30 minutes to go through it.
Your suite7 500 €3 weeks Up to 10 benchmarks, plus whether they measure ten things or two, plus the analysis restricted to your leading models, plus a written statement a procurement team can use.
Continuous1 200 €/moongoing The audit runs in your CI on every change to the suite and fails the build when an item turns negative or the effective length drops. Quarterly review.
Platform licence18 000 €/yrby contract For an evaluation platform embedding the audit in its own product.

Refund clause. If the audit finds neither an item running backwards nor redundancy above 20%, you pay nothing.

This is not generosity. We genuinely do not know what is in your data — the redundancy we have measured on public benchmarks runs from about a third to almost all of the test, and yours could be none. The clause puts that uncertainty on the seller, where it belongs.

What we need from you

A table, in any of these shapes:

trial,item,correct
gpt-x,q001,1
gpt-x,q002,0
claude-y,q001,1

Or the raw logs from lm-evaluation-harness, Inspect, promptfoo or OpenAI Evals, which we read directly. Details →

Your harness already produces this. It cannot compute a score without knowing which items were right.

Where this does not work

Fewer than 30 respondents. Item statistics need respondents — models, runs, versions, whatever plays that part. Below 30 we will say so rather than produce a report; below 50 the screening list flags about one healthy item in seven.

experiments/DETECTION-2026-09/RESULT.md

Graded scores. If your items are scored on a scale rather than right or wrong, the standard tool refuses the file rather than rounding it — rounding a 1-to-5 rubric once reported a perfectly healthy instrument as measuring with nothing. There is a separate path that takes the scale range.

scripts/graded_items.py

Items unrelated to the rest. This is the weakest of the three detections and stays under 0.8 even at two hundred respondents. That threshold sits close to ordinary sampling noise and no amount of care moves it much. We report it as the weak signal it is.

To commission an audit, or to ask whether your data fits: contact@itemaudit.fr.