CANONICAL HISTORY
Learning from Delayed Rewards
Christopher John Cornish Hellaby Watkins submitted his University of Cambridge PhD thesis 'Learning from Delayed Rewards' in May 1989. Watkins's own account states that the thesis described a range of reinforcement-learning algorithms, including one-step Q-learning, and sketched a proof of its convergence.
Evidence / resource
This page preserves the public LINEAiGE record and its first-party source relationship.
Record identity
LINEAiGE IDq-learning-1989