Operations Research · 2014 · 19 citations · 34 references
Mathematical ProgrammingEngineeringGame TheoryRich ClassOperations ResearchStochastic GameUncertainty QuantificationManagementVariance BoundCombinatorial OptimizationDecision TheoryMechanism DesignSequential Decision MakingProbability TheoryComputer ScienceMarkov Decision ProcessExploration V ExploitationOptimal Total RewardStochastic OptimizationDecision Science
We identify a rich class of finite-horizon Markov decision problems (MDPs) for which the variance of the optimal total reward can be bounded by a simple linear function of its expected value. The class is characterized by three natural properties: reward nonnegativity and boundedness, existence of a do-nothing action, and optimal action monotonicity. These properties are commonly present and typically easy to check. Implications of the class properties and of the variance bound are illustrated by examples of MDPs from operations research, operations management, financial engineering, and combinatorial optimization.
34
Choice Reviews Online · 1992 · 1.5K citations
The variance of discounted Markov decision processes
Matthew J. Sobel · Journal of Applied Probability · 1982 · 213 citations