2017 · 131 citations · 46 references
Artificial IntelligenceEngineeringMachine LearningSequential LearningIntelligent SystemsRobot LearningImitation LearningVisual Semantic PlanningVisual ObservationsVision Language ModelAction Model LearningComputer ScienceWorld ModelDeep LearningSuccessor RepresentationsComputer VisionDeep Reinforcement LearningAi PlanningVisual ReasoningPlanningRobotics
A crucial capability of real-world intelligent agents is their ability to plan a sequence of actions to achieve their goals in the visual world. In this work, we address the problem of visual semantic planning: the task of predicting a sequence of actions from visual observations that transform a dynamic environment from an initial state to a goal state. Doing so entails knowledge about objects and their affordances, as well as actions and their preconditions and effects. We propose learning these through interacting with a visual and dynamic environment. Our proposed solution involves bootstrapping reinforcement learning with imitation learning. To ensure cross task generalization, we develop a deep predictive model based on successor representations. Our experimental results show near optimal results across a wide range of tasks in the challenging THOR environment.
46
Laurens van der Maaten, Geoffrey E. Hinton · Journal of Machine Learning Research · 2008 · 35.7K citations
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver et al. · Nature · 2015 · 28.8K citations
Artificial Intelligence, Engineering, Deep Reinforcement Learning +3
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J. Maddison et al. · Nature · 2016 · 15.5K citations