PAPER DIGEST
Most Influential ICML 2005 Paper · 2026-03 edition

Reinforcement Learning With Gaussian Processes

Yaakov Engel; Shie Mannor; Ron Meir

Venue
International Conference on Machine Learning (ICML) 2005
Recognition
Most Influential ICML 2005 Paper (Rank No. 15)
Edition
2026-03
Impact factor
6
Certificate ID
1fdfb8946b9f01f0

Abstract

Gaussian Process Temporal Difference (GPTD) learning offers a Bayesian solution to the policy evaluation problem of reinforcement learning. In this paper we extend the GPTD framework by addressing two pressing issues, which were not adequately treated in the original GPTD paper (Engel et al., 2003). The first is the issue of stochasticity in the state transitions, and the second is concerned with action selection and policy improvement. We present a new generative model for the value function, deduced from its relation with the discounted return. We derive a corresponding on-line algorithm for learning the posterior moments of the value Gaussian process. We also present a SARSA based extension of GPTD, termed GPSARSA, that allows the selection of actions and the gradual improvement of policies without requiring a world-model.

Download PDF certificate