2022 · 14 citations · 17 references
Testing is a promising way to gain trust in a learned action policy π, in particular if π is a neural network. A “bug” in this context constitutes undesirable or fatal policy behavior, e.g., satisfying a failure condition. But how do we distinguish whether such behavior is due to bad policy decisions, or whether it is actually unavoidable under the given circumstances? This requires knowledge about optimal solutions, which defeats the scalability of testing. Related problems occur in software testing when the correct program output is not known.
17
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver et al. · Nature · 2015 · 28.8K citations
Artificial Intelligence, Engineering, Deep Reinforcement Learning +3
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J. Maddison et al. · Nature · 2016 · 15.5K citations
Kexin Pei, Yinzhi Cao, Junfeng Yang et al. · 2017 · 1.2K citations · Full text
Artificial Intelligence, Convolutional Neural Network, Deepfake Detection +12