2018 · 19 citations · 21 references
Artificial IntelligenceEngineeringMachine LearningAdvanced ComputingHardware AlgorithmComputer ArchitectureBootstrapped HeadsMulti-agent LearningVanilla DdpgData ScienceRobot LearningParallel ComputingDouble Bootstrapped ArchitectureAutonomous LearningComputer EngineeringComputer ScienceDeep LearningFpga DesignExploration V ExploitationReward HackingDeep Reinforcement LearningParallel Programming
Deep Deterministic Policy Gradient (DDPG) algorithm has been successful for state-of-the-art performance in high-dimensional continuous control tasks. However, due to the complexity and randomness of the environment, DDPG tends to suffer from inefficient exploration and unstable training. In this work, we propose Self-Adaptive Double Bootstrapped DDPG (SOUP), an algorithm that extends DDPG to bootstrapped actor-critic architecture. SOUP improves the efficiency of exploration by multiple actor heads capturing more potential actions and multiple critic heads evaluating more reasonable Q-values collaboratively. The crux of double bootstrapped architecture is to tackle the fluctuations in performance, caused by multiple heads of spotty capacity varying throughout training. To alleviate the instability, a self-adaptive confidence mechanism is introduced to dynamically adjust the weights of bootstrapped heads and enhance the ensemble performance effectively and efficiently. We demonstrate that SOUP achieves faster learning by at least 45% while improving cumulative reward and stability substantially in comparison to vanilla DDPG on OpenAI Gym's MuJoCo environments.
21
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver et al. · Nature · 2015 · 28.8K citations
Artificial Intelligence, Engineering, Deep Reinforcement Learning +3
Mastering the game of Go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan et al. · Nature · 2017 · 9K citations