Skip to main content
NexoraLaunch home

Build log live

Benchmarks and Evals You Can Defend in Public

AI · 10 min read ·

Performance claims for AI and developer tools survive scrutiny when the test is described, repeatable and honest. How to design and report evaluations.

Illustration: A deep-violet chart with thin gridlines, three bars with error whiskers and a footnote strip reading 'n, date, version, method'

A bar chart with one very tall bar is the most persuasive and least trustworthy image in technology marketing. Everyone has seen it: a product, a rival and a gap. Everyone has learned to ask what the test was, who ran it and whether the rival was set up fairly. If the answer is "we do not say", the chart does harm.

Claims about performance are not optional for AI and developer tools; buyers expect them. This guide explains how to design evaluations and report them so that they survive the questions a careful reader will ask.

What a benchmark is for

A benchmark, in computing, is the act of running a set of tests or programs against a system in order to assess its relative performance. Benchmarks give shared yardsticks: how fast does this library parse a file, how accurately does this model classify these documents.

They are useful because they make comparison possible. They are dangerous because any single number leaves out most of what matters, and because a benchmark can be designed, deliberately or not, to favour a particular approach.

Think of a benchmark as a measurement of one thing under defined conditions, never as a verdict on quality.

Start with the question

Before running anything, write the question the evaluation should answer. Examples:

  • "Does the new version parse typical configuration files faster than the previous one, on the same machine?"
  • "How often does the assistant draft a reply that a support agent would send with little editing?"
  • "On the kinds of documents our customers upload, how often does the extractor find all the required fields?"

A question tied to what users actually do yields claims users care about. A generic benchmark may produce a bigger number that means less.

Build a test set that reflects reality

For tasks based on data, the test set is everything.

Representative. It should resemble what users give the product, in type, difficulty and mess. A test set of clean, short examples will flatter a system that struggles with real inputs.

Separate. Machine-learning practice divides data into training, validation and test sets. The test set is held back from development, used only to estimate how well the system will do on new data. If you tune your product against the test set, the score loses its meaning.

Labelled carefully. Decide what counts as correct, write it down and have more than one person label a sample to check agreement.

Documented. Describe how it was assembled, how big it is, what is in it and what is deliberately excluded.

Refreshed. Keep adding new examples, especially failures found in the field, so the test set does not go stale.

Beware of overfitting and contamination

Overfitting occurs when a system learns the details of its training data so closely that it performs worse on new data. A related problem for evaluation is tuning to the test: if you adjust your system repeatedly while looking at the test score, you gradually fit it to that test and the score becomes an overestimate.

For models trained on large amounts of data, there is a further risk: contamination, where examples from a public benchmark appear in the training data, so that the model has effectively seen the answers. If you use a public benchmark, say whether you have checked for this, and prefer your own private test set for the claims that matter most.

Protect your test set: keep it separate, limit who can see it and note every time it is used.

Make comparisons fair

If you compare against alternatives, fairness is the issue.

  • Same task, same data, same conditions. Same hardware, same version, same settings.
  • Give rivals a fair set-up. Use their recommended configuration, not a deliberately poor one.
  • State versions and dates. Software changes; a comparison against an old version is out of date.
  • Report where you lose. If a rival wins on one measure, say so.
  • Use named versions, not "a leading tool".
  • Check with them, where practical. Giving a competitor the chance to correct a set-up error is courteous and improves accuracy.

Advertising rules in many countries require comparisons to be fair, accurate and capable of substantiation. In the UK, the Advertising Standards Authority administers codes that apply to comparative claims; take advice before using named comparisons in marketing.

Report like a scientist

A defensible report has a standard shape.

  1. The question.
  2. The set-up: system versions, hardware, settings, data.
  3. The method: how the test was run, how scoring worked.
  4. The results: numbers with their units and the sample size.
  5. The uncertainty: how much the result might vary.
  6. The limits: what the test does not show.
  7. The date and who ran it.
  8. How to reproduce it.

Reproducibility, the ability of others to obtain consistent results using the same data and methods, is the heart of credible measurement. If you can share the scripts and, where possible, the data, do. If you cannot share the data because it is private, share a description and a smaller public sample.

Handle uncertainty honestly

Results vary. A model may score differently on different runs; a speed test varies with machine load. Report that.

  • Run more than once and report the range or the spread.
  • Report the sample size. A score of ninety out of a hundred is less certain than nine hundred out of a thousand.
  • Do not claim a win within the noise. If the difference is smaller than the variation, say the results are comparable.
  • Use plain ranges rather than false precision. "Between eighty-five and ninety percent" is better than "87.3 percent" from a small test.

Include human judgement when it matters

Some qualities, such as helpfulness or readability, cannot be fully captured by an automated score. Use human reviewers for a sample, with clear criteria, and report how many reviewers, how they agreed and how they were instructed. When automated scorers, including other models, are used to judge outputs, say so, and check them against human judgement on a sample.

Report failures

A good evaluation shows where the system fails. Include a short analysis of errors: which kinds of input cause the most, which are rare but severe and what you are doing about them. Developers and buyers trust a report that shows its warts.

Keep claims in proportion

Write the claim to match the test. "Parsed our two hundred sample files faster than version two on the same machine, by a median of about a third" is a claim you can stand behind. "The fastest parser available" is not supported by a test that covered two versions.

Avoid single-number summaries for complex behaviour. If you must give one, accompany it with the breakdown.

A worked example

A team releases a library for extracting tables from documents. Their first claim: "The most accurate extractor on the market." They rewrite after planning a proper evaluation.

The question: "On the kinds of invoices and statements our users upload, how often does the extractor return every row and column correctly?" They assemble four hundred documents, drawn with permission from user-supplied samples and anonymised, held separate from development. Two people label the correct output for a hundred of them to check agreement; they agree on all but three, which they resolve.

They run their version and two named alternatives, using each tool's documented default settings, on the same machine, three times each. They report: "All rows and columns correct: our extractor, about eighty-two percent; alternative A, about seventy-four; alternative B, about sixty-nine; ranges across runs under two points. Our extractor is worse than alternative A on scanned documents with handwriting. Test set: four hundred documents; versions and date listed; scripts published; the documents themselves are private." They add a note that the result may not generalise to other document types.

The claim is modest and exact. A sceptical reader can rerun the scripts on their own documents, which is the best evidence of honesty a benchmark can give.

Questions engineers ask

Can I use another model to grade outputs? Yes, as a way to scale, but verify it. Compare its judgements with human ratings on a sample, report the agreement and watch for systematic biases, such as favouring longer answers.

How often should I rerun evaluations? On every significant change, and on a schedule, because dependencies and models change underneath you. Keep results in a table with dates so that regressions are visible.

What if the public benchmark scores look great but users complain? Trust the users. The benchmark may not measure what they need. Add their failures to your private test set.

Should I publish negative results? When they affect users' decisions, yes. A short note that a change reduced accuracy on one kind of input, and what you did, builds trust.

How do I avoid cherry-picking? Decide the test and the reporting format in advance, write it down and report all results from that design, including the unflattering ones.

A reporting template

Question: one sentence. Systems and versions: list with dates. Data: description, size and source. Method: steps and scoring. Results: a table with units and sample sizes. Spread: range across runs. Where we lose: one or two lines. Limits: what this test does not show. Reproduce: link to scripts and instructions. Date and author. Filling in this template takes an hour and turns a vague claim into a document that readers can check.

Summary

A defensible benchmark starts with a question that matters to users, uses a representative, separate and documented test set, guards against overfitting and contamination, compares fairly with named versions and reports uncertainty and failures. Describe the set-up and method so that others can reproduce the result, keep claims proportional to the test and date everything. Honest evaluation is slower than a headline bar chart and far more durable.

Questions and answers

What is a benchmark?
A standard test or set of tests used to compare the performance of systems, such as how fast software runs or how accurately a model answers questions.
What is overfitting to a test?
When a system is tuned so closely to a particular test set that it performs well on it but poorly on new data, making the score misleading.
Can I use public benchmarks for my product claims?
Yes, with care. Public benchmarks may not reflect your users' tasks, and may be contaminated if test data leaked into training.
How big does a test set need to be?
Large enough that the difference you claim is larger than the uncertainty. Report the size and be cautious with small samples.

Sources

Ask a question