PubMed · 2008 · 61 citations · 41 references
Artificial IntelligenceBayesian Decision TheoryEngineeringIntelligent SystemsLimited ReinforcementReinforcement Learning (Computer Engineering)Data ScienceUncertainty QuantificationManagementPlanning DomainsRobot LearningApproximation ApproachDecision TheoryCognitive ScienceAction Model LearningSequential Decision MakingComputer SciencePomdp Model ParametersExploration V ExploitationMarkov Decision ProcessAi PlanningBayes Risk
Partially Observable Markov Decision Processes (POMDPs) have succeeded in planning domains that require balancing actions that increase an agent's knowledge and actions that increase an agent's reward. Unfortunately, most POMDPs are defined with a large number of parameters which are difficult to specify only from domain knowledge. In this paper, we present an approximation approach that allows us to treat the POMDP model parameters as additional hidden state in a "model-uncertainty" POMDP. Coupled with model-directed queries, our planner actively learns good policies. We demonstrate our approach on several POMDP problems.
41
Learning to Predict by the Methods of Temporal Differences
Richard S. Sutton · Machine Learning · 1988 · 3.9K citations · Full text
A survey of robot learning from demonstration
Brenna Argall, Sonia Chernova, Manuela Veloso et al. · Robotics and Autonomous Systems · 2008 · 3.2K citations