Skip to content

AI Glossary · Concepts

What is Benchmark Contamination?

Benchmark Contamination

Benchmark contamination occurs when test questions, answers, or related evaluation information enter training or become available during a test in ways that compromise its intended measurement. A high score may reflect prior exposure or answer leakage rather than the capability being tested.

· Updated · Chain of Thought

AI Evaluation & Reliability

  • 1 explainer
  • 2 sources

Training may include a public test question, or a browsing agent may find the benchmark’s answer online during a run. These routes require different checks. AI benchmarks depend on the test conditions: access to an answer key changes what a score measures.

Anthropic’s BrowseComp analysis documents an agent finding evaluation material at test time. That is distinct from simply memorizing answers during training. The case illustrates why an evaluation needs to state which tools and sources are available during the run.

Engineering example: keep a private set of fresh product tasks for the final comparison, separate from the examples used to tune prompts. Record model and dataset versions, inspect suspicious answer matches, and distinguish legitimate retrieval from access to the answer key. Contamination is a reason to examine a score’s meaning; a suspiciously strong result alone does not prove it happened.

Sources

Go deeper