Benchmarks
LLM Evaluation Harnesses: Reproducibility and Hidden Variables
A practical guide to evaluation harness reproducibility: identify hidden variables, report the full evaluated system, and label benchmark comparisons honestly.
A benchmark score is only as reproducible as the system and procedure that produced it. Two teams can name the same model and benchmark yet report different numbers because the prompt wrapper, dataset revision, retry policy, test environment, or scorer changed. An evaluation harness is the code and configuration that turns those choices into measurements. Treating the harness as part of the experiment makes results explainable, comparable, and safer to use in engineering decisions.
The score is an evaluated system, not a model name
A useful mental model is:
Model + data + prompt + decoding + tools/retrieval + budgets/retries + environment + scorer/judge + harness version = reported result
The harness decides what is sent, how many chances a system receives, how outputs are parsed, and what counts as success. That means a benchmark result is evidence about a particular evaluated system, not a permanent property of a marketing model family. AI Model Benchmarks Explained provides the broader literacy; this article focuses on the machinery that makes one run reproducible.
A good harness makes every consequential choice explicit. It also preserves enough artifacts to explain a surprising result without asking an engineer to remember what changed between runs.
Pin model identity and provider state
Record the exact provider and model identifier, revision or checkpoint, endpoint version, retrieval date, quantization, fine-tune, adapter, and serving configuration when they affect behavior. A family label such as “model X” is not enough if an API silently routes to a newer revision or a local deployment uses a different tokenizer. For hosted services, exact bit-for-bit reproduction may be impossible when weights, safety layers, or infrastructure change. State that limitation and retain the date and endpoint information so an independent team can perform a meaningful controlled replication.
Keep generation settings beside the model identifier. Temperature, top-p, top-k, maximum output tokens, sampling mode, and the number of generated samples can all change the estimate. A seed is useful when supported, but temperature zero is not a universal proof of mathematical determinism: provider implementations, hardware, batching, and hidden service changes can still vary.
Identify the dataset, not just its name
Pin the benchmark release, split or subset, commit or hash where available, preprocessing, filtering, and excluded cases. “The benchmark” can refer to a refreshed repository, a different test split, or a locally cleaned copy. A private evaluation set should carry an immutable version and a change log; see Designing a Private LLM Evaluation Dataset for sampling and slice governance.
Store the task order and sampling policy when those choices can influence prompts, cache state, or few-shot selection. If examples are filtered after a run starts, preserve both the original manifest and the final eligible denominator. Never remove failures silently just to make an aggregate look cleaner.
Prompts and chat templates are hidden variables
For every task, preserve the system instruction, user wrapper, role formatting, few-shot examples, task-specific transformations, stop sequences, and chat-template version. An instruction model may receive materially different tokens depending on how its roles are serialized. There is no universal “correct” chat template; use the format required by the model and document the implementation used by the harness.
A benchmark paper that used one prompt cannot be compared directly with a run that added a system message, changed few-shot examples, or extracted only the first line of an answer. Keep prompt templates in version control, hash them in the run manifest, and review changes like code changes.
Context, retrieval, and tools
Record the maximum context, tokenizer or encoding, truncation strategy, document ordering, and few-shot placement. Silent truncation can remove the task specification or evidence and change difficulty. If the evaluation uses retrieval, also record the corpus snapshot, embedding model, index revision, top-k, reranker, query rewriting, and context assembly. Otherwise the result mixes model quality with retrieval quality.
Tools make the evaluated system even more specific. Capture the tool set, schema and version, permission scope, timeout, error behavior, network policy, and whether tool results are live, mocked, or replayed. A model with repository search and test execution is not the same system as a code-only completion model. Tool access should be reproducible and least-privileged; it is part of the harness, not an informal convenience.
Budgets and retries change the question
An evaluation can impose token, time, compute, turn, or tool-call budgets. Agents and coding systems often receive a bounded number of planning steps, test runs, or sub-agent calls. Report those limits and whether they are hard stops or soft warnings.
Retries are equally important. Three attempts with repair feedback are not equivalent to one attempt if the report records only final success. Record the retry trigger, maximum attempts, whether the best attempt is selected, and whether failed attempts remain in the denominator. If a benchmark uses pass@k, label the sampling budget: pass@1 and pass@10 answer different operational questions and carry different inference costs.
Scoring code and judges are measurement instruments
Pin the parser, normalization rules, matching logic, test suite, timeout, aggregation, and failure handling. A scorer revision can change a result without changing the model. Count invalid outputs, parser failures, timeouts, execution failures, refusals, and API errors explicitly.
If an LLM judges quality, record its model and revision, prompt and rubric, ordering, temperature, number of judgments, aggregation, and tie handling. Human ratings need a rubric, annotator instructions, sampling or blinding choices where relevant, adjudication, and agreement measures. A judge is not an oracle; it is another variable in the measurement system. How to Compare AI Models for a Production Use Case shows how to connect these measurements to a decision without hiding uncertainty.
Environment, caches, and concurrency
Executable tasks depend on the operating system, runtime, compiler or interpreter, package lock, container image, dependencies, hardware, resource limits, and network policy. Preserve an image digest or environment lock where feasible. Generated code is untrusted: run it with isolation, timeouts, restricted filesystem and network access, and bounded processes. The official HumanEval harness itself warns that model-generated code must be run inside a robust sandbox; a benchmark result is not a reason to weaken execution controls.
State whether caches are enabled. Caches can change latency, cost, provider response reuse, retrieval freshness, and even correctness when stale tool state is reused. For performance runs, record parallelism, warm-up, queueing, and provider rate limits. Otherwise one team may be measuring a warm cache while another is measuring throttled, cold execution.
Hidden variable matrix
| Variable | Example | How it changes the result | Must record? | Comparison risk |
|---|---|---|---|---|
| Model revision | API snapshot or checkpoint | Changes capability and refusals | Yes | Family-name comparisons mislead |
| Dataset revision | Split, filter, commit | Changes task difficulty and denominator | Yes | Same benchmark name, different cases |
| Prompt/template | System message, chat wrapper | Changes instructions and token budget | Yes | Scores are not prompt-equivalent |
| Decoding | Temperature, samples, seed | Changes variance and pass@k | Yes | More search looks like better quality |
| Context | Truncation and ordering | Removes evidence or requirements | Yes | Hidden truncation changes difficulty |
| Tools/retrieval | Search, index, top-k | Adds capabilities and external quality | Yes | Model and scaffold are conflated |
| Retry/budget | Repair attempts, step limit | Gives extra chances to recover | Yes | Final-only reporting hides failures |
| Scorer/judge | Tests, rubric, judge model | Defines what counts as success | Yes | Measurement changes unnoticed |
| Environment | Image, dependencies, timeout | Alters execution and failures | Yes | “Same code” is not same runtime |
| Cache/concurrency | Reuse, parallel workers | Alters latency, freshness, throttling | When material | Performance claims become incomparable |
| Failure policy | Drop, retry, or count error | Changes denominator and aggregate | Yes | Reliability is overstated |
Produce a reproducibility manifest
Every run should emit a machine-readable manifest alongside results. At minimum include:
- run identifier and timestamp;
- model, revision, provider, and endpoint;
- dataset, split, filters, and revision hash;
- harness repository and commit;
- prompt, chat-template, and decoding configuration;
- context and truncation policy;
- tool schemas, permissions, retrieval and index configuration;
- token, time, step, tool, concurrency, and retry budgets;
- scorer or judge versions and aggregation rules;
- environment or container digest, dependencies, and resource limits;
- seed and cache policy where applicable;
- sample count, failure counts, result file hashes, and artifact locations.
Preserve per-case outputs, scores, errors, logs, and aggregates subject to privacy and security controls. Hashing artifacts helps teams prove that a reviewed run is the one being discussed. Do not log secrets or unrestricted private prompts merely for reproducibility; redact sensitive fields while retaining stable references.
Comparability levels
PermsAI uses this operational taxonomy:
- Level A — identical-run reproduction. The same dataset and artifacts, harness commit, configuration, environment, and where possible provider state are used again. Small hosted-service differences should still be disclosed.
- Level B — controlled replication. An independent team follows the same methodology and pinned versions but may use a separate machine or provider account. Differences are investigated through the manifest.
- Level C — nominal benchmark comparison. Teams share a benchmark name but lack evidence that prompts, budgets, tools, scoring, and environment match. Treat the result as directional, not a head-to-head proof.
These levels are PermsAI guidance, not an official NIST or HELM standard. Labeling the level prevents a polished leaderboard number from being mistaken for a reproducible experiment.
A practical harness workflow
- Freeze inputs. Create immutable dataset and model references, prompt files, scorer code, and an environment lock.
- Declare policy. Set decoding, context, tools, budgets, retries, timeouts, concurrency, cache, and failure accounting before running.
- Run a calibration slice. Inspect representative cases, parser behavior, and sandbox failures before spending the full budget.
- Execute and log. Emit the manifest, per-case records, telemetry, and hashes in the same run directory.
- Review denominator and slices. Report eligible cases, failures, abstentions, and important subgroups, not just one aggregate.
- Re-run changed components. If a prompt, scorer, dependency, or provider revision changes, start a new run ID and explain the delta.
- Publish limitations. State what cannot be reproduced exactly and what evidence supports the comparison level.
Coding and benchmark claims checklist
Before repeating a score, ask: Which benchmark and version? Which model revision? Which prompt and scaffold? Which tools and retrieval? What sample, retry, and test-feedback budget? Which container, dependencies, timeout, and resource limits? Which parser, tests, scorer, or judge? How many cases failed, timed out, or produced invalid output? Are results cached? Are artifacts and hashes available?
A harness that answers these questions turns a number into an auditable engineering result. The goal is not to promise impossible determinism. The goal is to make differences explainable, replications meaningful, and model comparisons honest.