CANONICAL HISTORY
Q-learning
Christopher J. C. H. Watkins and Peter Dayan published Q-learning in Machine Learning in May 1992. The paper describes Q-learning as an incremental dynamic-programming method for agents to improve action-value estimates and presents a detailed convergence theorem for the method under stated conditions.
Evidence / resource
This page preserves the public LINEAiGE record and its first-party source relationship.
Record identity
LINEAiGE IDq-learning-1992