Benchmarks
Coding Benchmarks Explained: HumanEval, SWE-bench, and Beyond
A practical guide to interpreting coding AI benchmarks: compare function completion, repository repair, agent scaffolds, test quality, and reproducibility.
Coding benchmarks are useful only when you know what they actually measure. HumanEval, SWE-bench, and newer coding evaluations all put a model in front of software-related tasks, but they differ in task design, execution harness, success criteria, and the amount of scaffolding around the model. A score is therefore evidence about an evaluated system, not a permanent ranking of a model name. AI Model Benchmarks Explained covers general benchmark literacy; this guide focuses on coding-specific interpretation.
Start with the task, not the leaderboard
The first question is what the evaluated system is asked to produce. A function-completion benchmark supplies a relatively bounded specification and checks whether generated code satisfies tests. A repository-repair benchmark supplies a codebase, issue description, and project context, then checks whether a proposed patch resolves the issue. An agentic coding task can add shell tools, search, test execution, planning turns, and recovery steps. Those are different capabilities even when every result is reported as a percentage.
Task wording also determines what counts as available information. Is the model given only a function signature, a natural-language issue, failing tests, documentation, or an entire repository? Are hidden tests used? Can the system inspect dependencies or internet resources? A benchmark name does not answer those questions. Read the task card, repository, and evaluation configuration before comparing scores.
What HumanEval measures
The official OpenAI HumanEval repository describes a hand-written evaluation set for code-generation problem solving. Each example asks a model to complete a function from a prompt containing a signature, docstring, and surrounding code context. The repository provides an evaluation harness and reports pass@k-style results: whether at least one of k sampled completions passes the supplied tests. This setup is valuable for measuring constrained function completion, but it is not a complete measure of software engineering.
The execution step is a security boundary. HumanEval's own documentation warns that generated code is untrusted and strongly encourages running it in a robust sandbox. A reproducible run should therefore document the language runtime, dependencies, test command, timeout, process limits, filesystem and network policy, and how failures are classified. Never treat a benchmark harness as permission to execute arbitrary model output on a developer workstation.
Pass@1 and pass@10 answer different questions. Pass@1 approximates one-shot success under the declared sampling procedure. Pass@10 gives the system more attempts and can be useful when a product is allowed to search, but it also consumes more tokens and compute. Report the sampling budget and preserve failed attempts; a final successful sample is not equivalent to one attempt that succeeded immediately.
HumanEval can reveal syntax, API-use, and local reasoning weaknesses. It is less informative about repository navigation, issue interpretation, dependency changes, review quality, long-running tool use, or maintaining a large codebase. Its compact tasks also make contamination and memorization important evaluation concerns. Use a pinned dataset revision and disclose any filtering or prompt wrapper.
What SWE-bench measures
The official SWE-bench project evaluates language models on real-world software issues collected from GitHub repositories. A system receives a repository state and an issue, generates a patch, and is evaluated against project tests and the benchmark's task-specific infrastructure. The benchmark is consequently closer to repository-level repair than isolated function completion. It exercises issue understanding, code search, edits across files, dependency and test awareness, and the ability to produce a patch that works in a real project context.
That realism introduces more variables. Repository checkout, commit, build tools, package versions, operating-system image, test selection, patch application, and timeout all affect outcomes. The project uses containerized evaluation to make environments more consistent, but containers are not a guarantee of identical behavior. Record the image or environment version, harness commit, resource limits, network policy, and whether tests were rerun after a repair attempt.
A SWE-bench score also depends on how the system is scaffolded. One system may use a single completion; another may browse files, run tests, ask for feedback, and retry. Tool access and iteration budgets can improve the result while increasing cost and latency. Compare systems only when their task set, context, tools, patch policy, and recovery budget are materially aligned. LLM Evaluation Harnesses: Reproducibility and Hidden Variables provides a manifest for recording those hidden variables.
HumanEval versus SWE-bench
| Dimension | HumanEval | SWE-bench | Interpretation risk |
|---|---|---|---|
| Primary task | Complete a specified function | Resolve an issue in a real repository | Function skill is not repository repair skill |
| Context | Prompt, signature, and docstring | Issue plus repository files and project state | More context changes both difficulty and opportunity |
| Success test | Task tests for generated function | Project-level tests and patch evaluation | Test coverage and flakiness affect scores |
| Typical output | Code completion | Multi-file patch | Patch quality includes integration effects |
| Tooling | Often limited execution harness | Commonly search, edit, test, and environment tools | Scaffold can dominate model-only comparisons |
| Sampling | pass@k is common | Often one or more repair attempts | Attempts and retries must be reported |
| Environment | Language runtime and sandbox | Repository-specific container/build environment | Dependencies and system state matter |
| Best use | Controlled code-generation capability | Repository-level repair and engineering workflow | Neither is a universal measure of coding quality |
The matrix is not a claim that one benchmark is better. It is a reminder to match the benchmark to the product question. If you need to choose a completion model for a narrow API, HumanEval-like tasks may be a useful component. If you need an assistant that fixes issues in an existing service, repository-level tasks and your private evaluations are more relevant. How to Compare AI Models for a Production Use Case shows how to connect that choice to representative quality, safety, latency, and cost measurements.
Model capability versus scaffold capability
Coding evaluations often measure a model together with a scaffold. The scaffold includes prompt templates, repository indexing, file selection, tool schemas, test feedback, patch extraction, retry logic, and stopping rules. A result should state whether the reported number is model-only, model plus a fixed scaffold, or an end-to-end agent.
This distinction prevents two common errors. First, attributing a search or test-run advantage entirely to the model. Second, assuming an improvement is robust when it came from a changed prompt, larger context, or extra retries. Keep model revision, harness commit, task revision, tool permissions, context limits, and budgets in a run manifest. Designing a Private LLM Evaluation Dataset explains how to version representative internal tasks without leaking proprietary code.
Test quality is part of the measurement
Passing tests are evidence, not a full definition of correctness. A weak or incomplete test suite can accept a patch that violates security, performance, compatibility, or maintainability requirements. Conversely, a flaky or environment-sensitive test can reject a correct patch. Report the test command, version, timeout, flaky-test policy, and whether manual review or additional static checks were used.
For production decisions, supplement public coding benchmarks with tests that represent your failure modes: authorization boundaries, tenant ownership, secret handling, dependency updates, malformed input, migrations, observability, and rollback behavior. Include negative cases and hidden holdout tasks. Keep a denominator that distinguishes pass, fail, timeout, invalid patch, infrastructure error, and abstention rather than collapsing every non-pass into one unexplained number.
Contamination and task leakage
Public coding tasks can appear in training data, prompt examples, agent memories, or tool indexes. A high score may reflect memorization, benchmark-specific tuning, or a scaffold that recognizes known task structure. It does not prove broad generalization. Track task provenance, release dates, repository commits, and any filtering used to remove known overlaps. Rotate private holdouts and keep their access controlled.
Contamination is a reason to interpret a score cautiously, not a reason to discard every public benchmark. Public tasks are useful for trend analysis when versions and procedures are explicit. Pair them with fresh, private, and adversarial slices, and disclose which evidence supports your claim.
A coding-benchmark reading workflow
- Define the engineering decision: completion, repair, review, migration, or agent workflow.
- Pin the benchmark version, task list, repository commits, and exclusions.
- Record model revision, prompt, context window, decoding, tools, and scaffold commit.
- Declare sampling, retry, test-feedback, token, time, and concurrency budgets before running.
- Execute generated code or patches in an isolated, resource-limited environment.
- Preserve per-task outputs, test logs, failure classes, and artifact hashes.
- Review slices such as language, repository size, issue type, and security sensitivity.
- Re-run changed components under a new run identifier and explain the delta.
- Compare against a simple baseline and your private task set.
- Label the conclusion as directional, controlled, or reproducible according to the available evidence.
Claims checklist
Before repeating a coding benchmark claim, ask:
- Which exact task release and repository commits were evaluated?
- Which model revision, prompt, and chat template were used?
- Was the result one-shot, pass@k, or an iterative agent run?
- Which tools, context limits, retries, and test feedback were allowed?
- What runtime, container, dependencies, network policy, and resource limits applied?
- How were flaky tests, invalid patches, timeouts, and infrastructure failures counted?
- Is there evidence of contamination or benchmark-specific tuning?
- Can another team access the manifest, logs, and hashes needed to reproduce it?