2018 · 123 citations · 32 references
Artificial IntelligenceEngineeringMachine LearningHuman Pose EstimationIntelligent RoboticsEvaluation MetricsCognitive RoboticsIntelligent SystemsPredict Human MotionOpenpose LibraryRecent Deep LearningGlobal DiscriminatorMotion PredictionHuman MotionRobot LearningKinematicsHealth SciencesMachine VisionMotion SynthesisComputer ScienceHuman Image SynthesisDeep LearningComputer VisionGenerative Adversarial NetworkVideo HallucinationRobotics
Teaching a robot to predict and mimic human motion from historical movements is a crucial first step in human‑robot interaction, yet existing forecasting algorithms suffer from error accumulation and inaccurate prediction due to a lack of high‑level fidelity validation. The paper instruments a robot with prediction ability by leveraging recent deep learning and computer vision techniques. The system uses the robot camera to produce a human skeleton via OpenPose, conditions on the historical sequence, and forecasts plausible motion with a motion predictor, while a global discriminator inspired by GANs ensures smooth, realistic predictions. The motion GAN model outperforms state‑of‑the‑art approaches on the H3.6M dataset and enables the robot to replay predicted motion in a human‑like manner during interaction.
Teaching a robot to predict and mimic how a human moves or acts in the near future by observing a series of historical human movements is a crucial first step in human-robot interaction and collaboration. In this paper, we instrument a robot with such a prediction ability by leveraging recent deep learning and computer vision techniques. First, our system takes images from the robot camera as input to produce the corresponding human skeleton based on real-time human pose estimation obtained with the OpenPose library. Then, conditioning on this historical sequence, the robot forecasts plausible motion through a motion predictor, generating a corresponding demonstration. Because of a lack of high-level fidelity validation, existing forecasting algorithms suffer from error accumulation and inaccurate prediction. Inspired by generative adversarial networks (GANs), we introduce a global discriminator that examines whether the predicted sequence is smooth and realistic. Our resulting motion GAN model achieves superior prediction performance to state-of-the-art approaches when evaluated on the standard H3.6M dataset. Based on this motion GAN model, the robot demonstrates its ability to replay the predicted motion in a human-like manner when interacting with a person.
32
Sepp Hochreiter, Jürgen Schmidhuber · Neural Computation · 1997 · 93.8K citations
Unpaired Image-to-Image Translation Using Cycle-Consistent Adversarial Networks
Jun-Yan Zhu, Taesung Park, Phillip Isola et al. · 2017 · 21.3K citations · Full text
Engineering, Machine Learning, Image-to-image Translation +17
Automatic differentiation in PyTorch
Adam Paszke, Sam Gross, Soumith Chintala et al. · 2017 · 11.1K citations