Agentic Reinforcement Learning
Agentic reinforcement learning trains a model to choose actions across an agent's task, using rewards from interactions with tools or an environment.
Also known as: agentic RL
A training run can include several decisions: search for evidence, inspect a result, call another tool, and produce an answer. Rewards from those interactions guide updates to the model. Microsoft’s Agent Lightning illustrates this separation: an agent executes tasks, its traces and rewards become training data, and a training system updates the model for later runs. Training does not require exposing a deployed production application to unrestricted actions.
In episode 69 at 15:29, Joel Hron describes Thomson Reuters training its model with access to Westlaw and Practical Law. He connects this approach to the need to retrieve changing legal information rather than relying on knowledge stored during pre-training.
Connecting a search API enables tool use; it does not by itself train a better search policy. For a builder considering reinforcement learning, a useful first question is what counts as task success. Rewarding a completed tool call alone can miss whether the agent retrieved relevant evidence and answered correctly.
Sources
- Microsoft Research: Agent Lightning — Describes connecting agent execution to reinforcement learning for tool use and multi-step tasks.
Go deeper
- Luo et al.: Agent Lightning paper
Study how agent execution traces, rewards and credit assignment connect to reinforcement learning.