AI Glossary

Temporal difference learning

Temporal-difference (TD) learning updates predictions using a difference between successive estimates. In reinforcement learning, TD(0) moves a state’s value toward the observed reward plus the discounted estimate of the next state’s value.

Also known as: TD learning

· Updated · Chain of Thought

Suppose a robot estimates its current state’s value as 4. It moves, receives reward 1 and estimates the next state’s value as 6. With discount factor 0.9, the target is 1 + 0.9 × 6 = 6.4. A learning rate of 0.5 moves the old estimate halfway toward that target: 4 + 0.5 × (6.4 − 4) = 5.2. These are illustrative settings.

The next-state value is an estimate, not an observed final outcome. Using it in the target is called bootstrapping. TD can update after a transition, while Monte Carlo prediction waits for the sampled return from a completed episode. For a true terminal state, the next-state value is zero.

This helps with long or continuing tasks, but an inaccurate value estimate also affects the target. Q-learning applies this idea to action values and uses the best estimated next action. Other TD methods use different targets.

Sources

Go deeper