Reinforcement Learning with Verifiable Rewards
Reinforcement learning with verifiable rewards (RLVR) trains a model using rewards from mechanically checkable outcomes, such as a math answer checker or code tests. The checks determine which outputs receive reinforcement.
Also known as: RLVR
Model ArchitectureAI Evaluation & Reliability
A math task can reward a matching final answer; a coding task can reward passing predefined tests. DeepSeek-R1 describes rule-based accuracy rewards for such tasks, along with a separate reward for response format. Checkable outcomes can supply a training signal without a human rating every response.
For example, a model asked to implement addition could earn a reward from tests containing only “2 + 2.” Hardcoding 4 passes that check while failing other inputs. The reward is verifiable, but its coverage is too narrow. This is a practical reward-hacking concern, not evidence that every test-based training run produces that behavior.
RLVR names a reinforcement-learning setup, not a proof of safe or general reasoning. Keep evaluation cases outside the reward checks and test constraints the verifier does not cover. It also differs from spending more test-time compute on a deployed request; the former changes model weights through training.
Sources
- DeepSeek-AI: DeepSeek-R1 — Section 2.2.2 describes rule-based accuracy and format rewards for reasoning training.
Go deeper
- DeepSeek-AI: DeepSeek-R1 repository docs
Inspect model releases, usage recommendations and local inference examples separate from the training paper.