LINEAiGE
CANONICAL HISTORY

Learning from Delayed Rewards

Christopher John Cornish Hellaby Watkins submitted his University of Cambridge PhD thesis 'Learning from Delayed Rewards' in May 1989. Watkins's own account states that the thesis described a range of reinforcement-learning algorithms, including one-step Q-learning, and sketched a proof of its convergence.

May 1989CANONICAL HISTORY

Evidence / resource

This page preserves the public LINEAiGE record and its first-party source relationship.

Record identity

LINEAiGE ID
q-learning-1989