AI Glossary

Q-learning

Q-learning is a model-free reinforcement-learning algorithm that learns values for state-action pairs. Its target uses the reward plus the discounted maximum estimated value of a next-state action, even when the agent’s behavior explores other actions.

· Updated · Chain of Thought

A robot in a small grid could store a table of Q-values, one for each state and available action. Suppose the current action value is 2, the received reward is 1 and the next state’s action values are 3 and 5. With discount 0.9, the target is 1 + 0.9 × 5 = 5.5. A learning rate of 0.5 gives 2 + 0.5 × (5.5 − 2) = 3.75. These are illustrative settings.

The robot can explore by taking a lower-valued next action while the update still targets the maximum. That is the off-policy distinction. SARSA instead uses the value of the next action actually selected. For a true terminal state, there is no future action value to add.

Q-learning is a temporal-difference control method. The tabular convergence result requires repeated exploration and suitable learning rates; it does not automatically extend to neural approximations. A large state space also makes an explicit table impractical.

Update toward the best next action valueThe old Q-value is 2, reward is 1 and the next-state action values are 3 and 5. With discount 0.9 the target is 1 plus 0.9 times 5, or 5.5. Learning rate 0.5 updates the current Q-value to 3.75. The max target is used even if behavior chooses the lower-valued next action. Update toward the best next action valueAn illustrative nonterminal transition with two available next actions Current Q(s, a) = 2Observed reward r = 1Next state: Q(s′, left) = 3Next state: Q(s′, right) = 5Target = r + γ max Q(s′, a′)Target = 1 + 0.9 × max(3, 5) = 5.5New Q = 2 + 0.5 × (5.5 − 2) = 3.75Behavior may explore “left”; the learning target still uses max(3, 5).For a true terminal state, use target = reward, with no future-value term.
A worked application of Watkins and Dayan’s Q-learning update. Values, discount and learning rate are illustrative. A true terminal transition has no future-action term. Download the image

Sources

  • Watkins and Dayan: Q-learning — Defines the maximum-next-action backup and proves tabular convergence under exploration and learning-rate conditions.

Go deeper