arXiv (Cornell University) · 2021 · 26 citations · 27 references
Using a high Update-To-Data (UTD) ratio, model-based methods have recently achieved much higher sample efficiency than previous model-free methods for continuous-action DRL benchmarks. In this paper, we introduce a simple model-free algorithm, Randomized Ensembled Double Q-Learning (REDQ), and show that its performance is just as good as, if not better than, a state-of-the-art model-based algorithm for the MuJoCo benchmark. Moreover, REDQ can achieve this performance using fewer parameters than the model-based method, and with less wall-clock run time. REDQ has three carefully integrated ingredients which allow it to achieve its high performance: (i) a UTD ratio >> 1; (ii) an ensemble of Q functions; (iii) in-target minimization across a random subset of Q functions from the ensemble. Through carefully designed experiments, we provide a detailed analysis of REDQ and related model-free algorithms. To our knowledge, REDQ is the first successful model-free DRL algorithm for continuous-action spaces using a UTD ratio >> 1.
27
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver et al. · Nature · 2015 · 28.8K citations
Artificial Intelligence, Engineering, Deep Reinforcement Learning +3
Continuous control with deep reinforcement learning
Timothy Lillicrap, Jonathan J. Hunt, Alexander Pritzel et al. · arXiv (Cornell University) · 2015 · 5.4K citations · Full text
Playing Atari with Deep Reinforcement Learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver et al. · arXiv (Cornell University) · 2013 · 5.1K citations · Full text
Artificial Intelligence, Convolutional Neural Network, Reward Hacking +11
MuJoCo: A physics engine for model-based control
Emanuel Todorov, Tom Erez, Yuval Tassa · 2012 · 4.3K citations