Temporal difference learning
Temporal-difference (TD) learning updates predictions using a difference between successive estimates. In reinforcement learning, TD(0) moves a state’s value toward the observed reward plus the discounted estimate of the next state’s value.
Also known as: TD learning
Suppose a robot estimates its current state’s value as 4. It moves, receives reward 1 and estimates the next state’s value as 6. With discount factor 0.9, the target is 1 + 0.9 × 6 = 6.4. A learning rate of 0.5 moves the old estimate halfway toward that target: 4 + 0.5 × (6.4 − 4) = 5.2. These are illustrative settings.
The next-state value is an estimate, not an observed final outcome. Using it in the target is called bootstrapping. TD can update after a transition, while Monte Carlo prediction waits for the sampled return from a completed episode. For a true terminal state, the next-state value is zero.
This helps with long or continuing tasks, but an inaccurate value estimate also affects the target. Q-learning applies this idea to action values and uses the best estimated next action. Other TD methods use different targets.
Sources
- Sutton and Barto: Temporal Difference Learning slides — Presents TD(0), Monte Carlo targets, terminal values and TD control algorithms.
Go deeper
- Sutton: Learning to Predict by the Methods of Temporal Differences paper
Read the original prediction-learning formulation and its comparison with outcome-based learning.