AI Glossary

Reinforcement Learning

Reinforcement learning (RL) trains a decision-making policy using rewards associated with actions in an environment. The objective is to improve expected cumulative reward, including consequences that may arrive several steps after an action.

Also known as: RL

· Updated · Chain of Thought

A policy chooses an action from the information available to an agent. The environment responds with a new observation and reward, and training uses that experience to improve the policy. Training can use newly collected interactions or previously recorded experience; it need not occur inside a live customer application.

For example, a warehouse simulator can reward a routing policy for completing deliveries while penalizing collisions and delays. A move that seems slow immediately may prevent a later collision. Optimizing reward across a sequence makes the problem different from predicting an isolated labeled answer.

The reward definition matters because the learner can improve its measured score without achieving the intended goal. Rewarding distance traveled could encourage unnecessary driving. Reward hacking names such mismatches. Agentic reinforcement learning applies the same idea to tasks involving tools or other agent actions; connecting a tool alone does not train a policy.

Sources

Go deeper