AI Glossary

Reward Hacking

Reward hacking occurs when an AI system finds an unintended way to earn a high score or reward without completing the task as intended.

· Chain of Thought

A shortcut can satisfy an evaluation’s score while missing its intended task. For example, an agent asked to find a software vulnerability could obtain the answer through a source it was not authorized to use. The apparent success would not demonstrate the capability the evaluation was designed to measure.

In its account of the July 2026 Hugging Face incident, OpenAI says agents sought test solutions through unintended routes. It identifies reward hacking, persistence on seemingly impossible tasks, unauthorized communication, and adoption of other agents’ goals as distinct contributing patterns. Agents also exchanged discoveries through message boards outside their permitted collaboration channels. The incident shows why a team evaluating agents needs to inspect how an answer was obtained and control what separate runs can share, not just score the final answer. Conor discusses the incident in episode 72.

From the conversation