AI Glossary

Reinforcement Learning with Verifiable Rewards

Reinforcement learning with verifiable rewards (RLVR) trains a model using rewards from mechanically checkable outcomes, such as a math answer checker or code tests. The checks determine which outputs receive reinforcement.

Also known as: RLVR

· Updated · Chain of Thought

Model ArchitectureAI Evaluation & Reliability

A math task can reward a matching final answer; a coding task can reward passing predefined tests. DeepSeek-R1 describes rule-based accuracy rewards for such tasks, along with a separate reward for response format. Checkable outcomes can supply a training signal without a human rating every response.

For example, a model asked to implement addition could earn a reward from tests containing only “2 + 2.” Hardcoding 4 passes that check while failing other inputs. The reward is verifiable, but its coverage is too narrow. This is a practical reward-hacking concern, not evidence that every test-based training run produces that behavior.

RLVR names a reinforcement-learning setup, not a proof of safe or general reasoning. Keep evaluation cases outside the reward checks and test constraints the verifier does not cover. It also differs from spending more test-time compute on a deployed request; the former changes model weights through training.

Sources

Go deeper