Training may include a public test question, or a browsing agent may find the benchmark’s answer online during a run. These routes require different checks. AI benchmarks depend on the test conditions: access to an answer key changes what a score measures.
Anthropic’s BrowseComp analysis documents an agent finding evaluation material at test time. That is distinct from simply memorizing answers during training. The case illustrates why an evaluation needs to state which tools and sources are available during the run.
Engineering example: keep a private set of fresh product tasks for the final comparison, separate from the examples used to tune prompts. Record model and dataset versions, inspect suspicious answer matches, and distinguish legitimate retrieval from access to the answer key. Contamination is a reason to examine a score’s meaning; a suspiciously strong result alone does not prove it happened.
Sources
- Sainz et al.: NLP Evaluation in trouble — Explains training on benchmark test data and the resulting measurement problem; abstract checked October 7, 2026.
- Anthropic: Eval awareness in BrowseComp — Primary account of test-time access to leaked answers and an answer key; checked October 7, 2026.