Benchmarks

Test-Time Compute and Reasoning Benchmarks

A practical framework for reporting test-time compute, samples, search, tools, retries, quality, latency, and cost in reasoning benchmarks without hiding the budget.

By PermsAI Editorial Team
Test-Time Compute and Reasoning Benchmarks featured image

A benchmark score is a measurement of a configured system, not a permanent property of a model name. If one system is allowed more samples, longer generation, search, verification, tools, or retries, it may achieve a higher score by spending more work at test time. That can be the right production trade-off, but the comparison is incomplete unless the budget and selection method are reported. Test-time compute benchmarks make this resource allocation visible without requiring private chain-of-thought traces or unverifiable estimates of hidden hardware work.

For baseline benchmark vocabulary, see AI Model Benchmarks Explained.

The unit being compared

Define the evaluated system as the model revision plus its inference strategy, prompt, sampler, tools, verifier, retry policy, scorer, and harness. The same base model can have distinct operating points: one-shot generation, multiple independent samples, a search loop, or an agent with tests and repair. These are legitimate systems, but they answer different questions.

The first question should be: what decision will this benchmark inform? A research comparison may want equal attempt budgets. A production team may care about cost per successful task and tail latency. A regression suite may choose a fixed, low budget for fast feedback. How to Compare AI Models for a Production Use Case provides the broader selection frame; this page focuses on normalizing test-time work.

TEST-TIME RESOURCE EQUATION

Model

  • Samples
  • Generation or reasoning budget
  • Search
  • Verifier or selector
  • Tools
  • Retries = Evaluation result

The equation is deliberately observable. Record what the provider and harness expose, and mark provider-internal computation as unknown rather than inventing FLOPs or hidden reasoning-token counts.

Training compute versus test-time compute

Training compute is spent creating or post-training a model. Test-time or inference compute is spent solving each evaluation task. A smaller model with ten attempts and a verifier and a larger model with one attempt are different allocations of resources. Neither is intrinsically superior without an objective such as quality at a fixed cost, latency, or energy budget.

Do not use a training cutoff, parameter count, or model tier as a substitute for the inference protocol. A benchmark report should identify the model revision and then describe the work performed for every task.

What counts as test-time compute?

Common mechanisms include a longer output budget, multiple independent samples, best-of-N selection, self-consistency voting, a learned or rule-based verifier, search over candidates, iterative critique and revision, code execution, retrieval, browser or calculator calls, and retries after invalid output. These mechanisms are not computationally equivalent. A verifier can be cheap and deterministic, or it can call another model; a tool call can add network latency and a separate bill. Report the mechanism at a level that is reproducible without exposing private reasoning traces.

One sample versus many

A one-shot or pass@1-like regime gives the model one opportunity. Allowing N attempts increases the opportunity to find a correct candidate, especially when attempts are meaningfully independent. The score must therefore include N and the sampling settings. Do not compare a one-sample result with a 100-sample result using only the final success percentage.

Coding evaluations make this visible through pass@1 and pass@k-style reporting. The HumanEval repository documents the task and evaluation context; the attempt budget is part of the claim, not a footnote.

Best-of-N and self-consistency

Best-of-N generates candidates and selects one with tests, a reward model, a judge, or a heuristic. The resulting quality depends on both the generator and selector. Self-consistency aggregates multiple sampled solutions, often by voting over answers. It may help on some tasks, but majority vote can repeat a shared mistake. Report sample count, aggregation, seed or sampling settings where available, and the selector's version.

Verifiers and search

A verifier may check constraints, execute tests, score an answer, or ask another model to judge it. A search procedure may expand candidate plans or programs until a termination rule is met. A higher score can reflect a better verifier or search policy rather than a stronger base model. Keep descriptions at the strategy level: search depth or expansion budget, candidate count, verifier identity, and termination condition. Do not require or reproduce hidden chain-of-thought content.

Tools, retries, and repair

Retrieval, code execution, browser access, calculators, external solvers, and test runners change the evaluated system. A closed-book model and a tool-enabled agent should have separate result rows. Likewise distinguish one-shot output, retry after a schema error, test-guided repair, and a full restart. Record maximum retries and whether failed attempts consume budget.

Observable versus hidden compute

Hosted services may expose a high-level setting such as reasoning effort, a maximum output token count, or a sampling parameter without revealing internal FLOPs, hidden reasoning tokens, or search traces. Record the setting name, value, provider, and revision; do not claim that “high” means the same work across providers. If an internal quantity is not published, write “unknown” and compare observable operating points instead.

Token counts are a useful proxy for generated work, but tokenizers and architectures differ. More visible tokens do not imply proportionally more FLOPs, and hidden computation may vary. Wall-clock latency and provider-reported cost are often more decision-relevant than a guessed compute conversion.

Score without budget is incomplete evidence

A maximum score hides the resource price of reaching it. Measure several operating points for each system: for example, one, four, and sixteen samples; short and long output caps; or bounded search depths. Plot quality against the chosen budget. The curve reveals marginal return and makes it possible to compare a cheap plateau with an expensive late gain.

SCORE-VS-BUDGET CURVE

For each operating point, record:

Budget → Quality score, uncertainty, latency, and cost

Use the same task set and scorer where possible. Confidence intervals and repeated runs show whether an apparent improvement exceeds run-to-run noise. Do not assert a universal scaling law; diminishing returns depend on model, task, selector, and budget definition.

Compute-matched and quality-matched comparisons

A compute-matched comparison fixes a resource definition before looking at results. The match could be samples, visible output tokens, wall-clock time, dollars, GPU time, tool calls, or a composite envelope. There is no universal normalization: choose the dimension that corresponds to the decision and disclose trade-offs.

A quality-matched comparison asks how much budget each system needs to reach a target quality or error rate. This is often more useful for production: compare cost, latency, and reliability at the same service-level objective rather than rewarding the system that spends the most.

COMPUTE-NORMALIZED COMPARISON MATRIX

SystemQuality scoreSamplesObservable token/output budgetTool callsRetriesLatencyCostSelector/verifierNotes
System AReport score and interval1Cap or unknown00Median and tailPer taskNoneOne-shot baseline
System BReport score and intervalNCap or unknown00Median and tailPer taskVote or judgeMulti-sample
System CReport score and intervalNCap or unknownBoundedBoundedMedian and tailInclude toolsTests/verifierTool-enabled

Do not fill unobservable values with guesses. A table with “unknown” is stronger evidence than false precision.

Pareto frontiers for production choices

Quality, latency, and cost usually trade off. Plot them together and identify dominated configurations. A system is dominated when another is at least as good on the relevant quality measure and no worse on the resource dimensions being optimized. Keep the frontier visible before collapsing it into a weighted score; an arbitrary weight can hide an important tail-latency or safety constraint.

PRODUCTION DECISION FRONTIER

Quality vs Latency vs Cost

For each frontier point, attach task slices, uncertainty, failure modes, and the test-time method. A high-quality point that breaches a latency or budget limit is not a production option until the limit changes.

TEST-TIME COMPUTE RESULT CARD

Retain one card per result:

FieldRecord
Model/versionExact provider and revision
Benchmark/versionDataset, split, and commit or hash
Test-time methodOne-shot, sampling, search, tools, or agent loop
Samples/budgetCount, output cap, search or effort setting
ToolsNames, versions, network/retrieval policy
Retry policyMaximum, triggers, and budget accounting
Selector/verifierIdentity, version, and decision rule
ScoreMetric, interval, and slice results
Latency/costPer task, successful task, and useful tail measures
EnvironmentHarness, dependencies, concurrency, scorer
UnknownsProvider-internal compute or unreported settings

The card makes a leaderboard row auditable and prevents a system-level result from being presented as a pure base-model capability.

Reliability, safety, and agentic systems

More test-time work can improve quality, but it can also increase tool exposure, retries, data access, and opportunities for unsafe actions. The direction is not universal. Record whether additional samples are isolated, whether verifiers can access sensitive context, and whether a search loop can call tools repeatedly. Agentic benchmarks cover these systems in more depth; Agentic Benchmarks: Measuring Tools, Planning, and Reliability is the adjacent guide.

For a safe comparison, apply the same permission and tool policy across operating points or report the policy change explicitly. A verifier that executes generated code, for example, needs the same containment assumptions as the main evaluation. A higher score under broader authority is not directly comparable to a restricted run.

Reproducibility and hidden variables

Capture model revision, dataset version, prompt/template, sampling settings, sample budget, search strategy, tool configuration, retry rules, selector, scorer, environment, concurrency, latency, and cost. LLM Evaluation Harnesses: Reproducibility and Hidden Variables explains why these details can change a result. Freeze configurations before comparing systems and rerun anchor cases after material changes.

A benchmark result describes performance at a stated budget. It does not isolate “the model’s reasoning ability” when selection, tools, or retries contribute. That is acceptable when the claim is system performance at that operating point; it is misleading when the claim says the base model is intrinsically better.

Leaderboards and reporting discipline

A leaderboard entry should expose inference budget, tool policy, scorer, and retries whenever the protocol permits. A higher score achieved with a much larger budget is not an apples-to-apples capability comparison. Do not accuse a leaderboard of hiding information without evidence; instead state which fields are reported, which are missing, and how that limits interpretation.

Report median and tail latency where relevant, cost per evaluated task and per successful task, and the rate of failed or invalid attempts. If provider prices change, preserve the measurement date and units rather than hard-coding a permanent dollar ranking.

Practical workflow

  1. Define the decision and the resource dimension that matters.
  2. Freeze model, data, prompt, scorer, and harness versions.
  3. Choose observable operating points and a retry/tool policy.
  4. Run the same slices at multiple budgets with uncertainty estimates.
  5. Record selector, verifier, tools, latency, and cost.
  6. Plot score-versus-budget and inspect diminishing returns.
  7. Compare compute-matched and quality-matched points.
  8. Mark Pareto-dominated configurations and preserve safety gates.
  9. Publish a result card with unknowns stated plainly.
  10. Revalidate after model, provider, harness, or tool changes.

Sources

These sources support test-time sampling, verification, search, coding-attempt budgets, and evaluation methodology. Specific comparisons still require the declared system protocol and observable budget; hidden provider computation remains unknown unless documented.