Run it yourself
Everything below runs on your machine and sends nothing anywhere. That is not a limitation we are dressing up: for a testing laboratory, a notified body or anyone under a confidentiality obligation, evaluation data cannot be uploaded to a stranger, and it does not have to be.
Python 3, no dependencies outside the standard library — except one, for recent Inspect logs, named below.
1. The free check, from a public leaderboard
Needs only what is already published: model names and their scores, plus how many items the test has.
python scripts/leaderboard_check.py --table leaderboard.csv --items 1000
It reports how wide the ruler is, how many models cannot be told apart from the leader, how many adjacent pairs are closer together than the measurement error, and how many of the orderings in your top twenty the scores do not support.
This is not the audit, and it says so on its own second page. An aggregate score cannot see a mis-keyed item, a dead item or the effective length — those need per-item data. A free check that lets you believe you have audited your benchmark has cost you more than it gave you.
It refuses rather than guesses. It will not run without the item count, because without it there is no sampling error and every claim it makes is built on one. And a board sitting entirely at or below 1.0 is refused rather than assumed: proportion and percentage differ by a factor of a hundred in the only quantity it computes, so a guess would not be a rounding, it would be the whole answer.
2. The real thing, from your harness logs
Four formats are read directly, so nobody has to write a converter before finding out whether they have a problem.
| Harness | What is read |
|---|---|
lm-evaluation-harness | the per-sample JSONL written by --log_samples |
Inspect | the .eval log, per-sample scores — read from its summaries, so a multi-gigabyte agent log is not loaded whole. Logs written by newer Inspect versions are zstd-compressed and need one package: pip install zstandard. Without it, the tool says so and reads nothing. |
promptfoo | the results JSON, one respondent per provider |
OpenAI Evals | the events JSONL |
python scripts/from_harness.py --out table.csv out/*/samples_*.jsonl
python scripts/item_analysis.py --table table.csv
Each conversion writes a provenance record with a checksum per source file, so the table can be traced back to the logs it came from.
One thing to check on your own logs. lm-evaluation-harness does not
write the model name into its sample files at all. The respondent name is derived from the directory
and the tool prints “derived from the directory name, not recorded in the log” every time it does
so. If you ran two models into one directory on different tasks, tell it the labels explicitly.
3. The plain table, from anything else
Three columns. Every harness on earth can emit this, because none of them can compute a score without it.
trial,item,correct
gpt-x,q001,1
gpt-x,q002,0
claude-y,q001,1
claude-y,q002,1
python scripts/item_analysis.py --table responses.csv
Column names are matched loosely: run_id, question_id,
is_correct and similar all work. Words are accepted as well as numbers —
true, pass, fail, no.
A graded score is refused, not rounded. If your correctness column holds a 1-to-5 rubric, the tool stops and points you at the graded path, which takes the scale range. Rounding it silently once turned a healthy instrument — alpha 0.864, seven of eight items carrying — into a report saying “a test of 8 items that measures with 0”.
4. The buyer-facing page, from a report
python scripts/instrument_report.py --report report.json --out report-dir/
Add your own costs and it states the saving in money and hours, with each multiplication printed beside its result:
python scripts/instrument_report.py --report report.json --out report-dir/ \
--cost-per-run 0.004 --runs-per-year 12 --seconds-per-run 9
The costing is quoted only on items the page is willing to recommend deleting, and it states in
its own words that it assumes every item costs the same to run — which is usually false, and which
--cost-table lets you replace with your real numbers.
Where to get it
The code is public. The statistics in it are a century old and public too.
If you would rather not read the output yourself, that is what the
audit is. If you run it and want to argue with the result, that is welcome:
write to contact@itemaudit.fr. Two of the corrections on the
evidence page came from someone objecting.