A coding model may generate several implementations and run each against tests. If any of k candidates passes, that trial counts as a success. On a code benchmark, pass@k estimates the fraction of tasks solved within k attempts. The HumanEval paper describes an estimator using n samples per task, where n is at least k, and the number that pass. This lets evaluators estimate pass@k from a larger sample pool rather than rely on one group of k attempts.
Example: a system that passes after ten attempts may be useful when those attempts are cheap and a trustworthy test suite selects the result. A high pass@k does not show that a deployment can identify the passing candidate without a verifier. The same score says little about a support bot expected to give one correct answer immediately.
Report k, the sampling setup, the verifier, and the computation budget. Do not compare pass@10 to another system’s pass@1 as though both had the same opportunity. Passing checks can also leave untested requirements unmet, so evaluation on your actual workload remains necessary.
Sources
- Chen et al.: Evaluating Large Language Models Trained on Code — Introduces the HumanEval functional-correctness evaluation and pass@k estimator.