2020 · 31 citations · 21 references
Artificial IntelligenceHuman PoseEngineeringIntelligent SystemsHuman-object InteractionNatural Language ProcessingMultimodal LlmVisual GroundingData ScienceVisual Question AnsweringHealth SciencesMachine VisionVision Language ModelComputer ScienceComputer VisionVisual ReasoningInteractive GraphObject RecognitionGraph-based Interactive ReasoningSemantic GraphActivity Recognition
Human-Object Interaction (HOI) detection devotes to learn how humans interact with surrounding objects via inferring triplets of < human, verb, object >. However, recent HOI detection methods mostly rely on additional annotations (e.g., human pose) and neglect powerful interactive reasoning beyond convolutions. In this paper, we present a novel graph-based interactive reasoning model called Interactive Graph (abbr. in-Graph) to infer HOIs, in which interactive semantics implied among visual targets are efficiently exploited. The proposed model consists of a project function that maps related targets from convolution space to a graph-based semantic space, a message passing process propagating semantics among all nodes and an update function transforming the reasoned nodes back to convolution space. Furthermore, we construct a new framework to assemble in-Graph models for detecting HOIs, namely in-GraphNet. Beyond inferring HOIs using instance features respectively, the framework dynamically parses pairwise interactive semantics among visual targets by integrating two-level in-Graphs, i.e., scene-wide and instance-wide in-Graphs. Our framework is end-to-end trainable and free from costly annotations like human pose. Extensive experiments show that our proposed framework outperforms existing HOI detection methods on both V-COCO and HICO-DET benchmarks and improves the baseline about 9.4% and 15% relatively, validating its efficacy in detecting HOIs.
21
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren et al. · 2016 · 214.9K citations · Full text
Image Classification, Deep Neural Networks, Machine Vision +14
Neural Message Passing for Quantum Chemistry
Justin Gilmer, Samuel S. Schoenholz, Patrick Riley et al. · arXiv (Cornell University) · 2017 · 3K citations · Full text
RMPE: Regional Multi-person Pose Estimation
Hao-Shu Fang, Shuqin Xie, Yu‐Wing Tai et al. · 2017 · 1.9K citations