Q-learning
Q-learning is a model-free reinforcement-learning algorithm that learns values for state-action pairs. Its target uses the reward plus the discounted maximum estimated value of a next-state action, even when the agent’s behavior explores other actions.
A robot in a small grid could store a table of Q-values, one for each state and available action. Suppose the current action value is 2, the received reward is 1 and the next state’s action values are 3 and 5. With discount 0.9, the target is 1 + 0.9 × 5 = 5.5. A learning rate of 0.5 gives 2 + 0.5 × (5.5 − 2) = 3.75. These are illustrative settings.
The robot can explore by taking a lower-valued next action while the update still targets the maximum. That is the off-policy distinction. SARSA instead uses the value of the next action actually selected. For a true terminal state, there is no future action value to add.
Q-learning is a temporal-difference control method. The tabular convergence result requires repeated exploration and suitable learning rates; it does not automatically extend to neural approximations. A large state space also makes an explicit table impractical.
Sources
- Watkins and Dayan: Q-learning — Defines the maximum-next-action backup and proves tabular convergence under exploration and learning-rate conditions.
Go deeper
- Hugging Face: Introducing Q-learning course
Follow the Q-table update and compare the behavior policy with its greedy learning target.