2001 · 288 citations · 14 references
This paper investigates the problem of policy learning in multiagent environments using the stochastic game framework, which we briefly overview. We introduce two properties as desirable for a learning agent when in the presence of other learning agents, namely rationality and convergence. We examine existing reinforcement learning algorithms according to these two properties and notice that they fail to simultaneously meet both criteria. We then contribute a new learning algorithm, WoLF policy hillclimbing, that is based on a simple principle: "learn quickly while losing, slowly while winning." The algorithm is proven to be rational and we present empirical results for a number of stochastic games showing the algorithm converges. 1
14
Andrew G. Barto · IFAC Proceedings Volumes · 1998 · 3K citations
Lloyd S. Shapley · Proceedings of the National Academy of Sciences · 1953 · 2.4K citations · Full text
Multiagent Systems: A Survey from a Machine Learning Perspective
Peter Stone, Manuela Veloso · Autonomous Robots · 2000 · 1.2K citations
The dynamics of reinforcement learning in cooperative multiagent systems
Caroline Claus, Craig Boutilier · 1998 · 1.1K citations