Reinforcement Learning
Reinforcement learning (RL) trains a decision-making policy using rewards associated with actions in an environment. The objective is to improve expected cumulative reward, including consequences that may arrive several steps after an action.
Also known as: RL
A policy chooses an action from the information available to an agent. The environment responds with a new observation and reward, and training uses that experience to improve the policy. Training can use newly collected interactions or previously recorded experience; it need not occur inside a live customer application.
For example, a warehouse simulator can reward a routing policy for completing deliveries while penalizing collisions and delays. A move that seems slow immediately may prevent a later collision. Optimizing reward across a sequence makes the problem different from predicting an isolated labeled answer.
The reward definition matters because the learner can improve its measured score without achieving the intended goal. Rewarding distance traveled could encourage unnecessary driving. Reward hacking names such mismatches. Agentic reinforcement learning applies the same idea to tasks involving tools or other agent actions; connecting a tool alone does not train a policy.
Sources
- OpenAI Spinning Up: Key concepts in RL — Defines policies, trajectories, environments, rewards and the expected-return objective.
Go deeper
- Hugging Face: Introduction to deep reinforcement learning course
Start with the agent-environment loop and work toward a trained game-playing agent.