IEEE Transactions on Circuits and Systems for Video Technology · 2022 · 26 citations · 32 references
Language GroundingToward Explainable 3DEngineeringMachine LearningVqa DatasetNew BenchmarkScene ModelingNatural Language ProcessingMultimodal LlmVisual GroundingComputational LinguisticsVisual Question AnsweringGrounded Question AnsweringLanguage StudiesCognitive ScienceMachine VisionVision Language ModelComputer ScienceDeep LearningComputer VisionStrong BaselineVisual ReasoningNew 3DExtended RealityLinguistics
Recently, 3D vision-and-language tasks have attracted increasing research interest. Compared to other vision-and-language tasks, the 3D visual question answering (VQA) task is less exploited and is more susceptible to language priors and co-reference ambiguity. Meanwhile, a couple of recently proposed 3D VQA datasets do not well support 3D VQA task due to their limited scale and annotation methods. In this work, we formally define and address a 3D grounded question answering (GQA) task by collecting a new 3D VQA dataset, referred to as flexible and explainable 3D GQA (FE-3DGQA), with diverse and relatively free-form question-answer pairs, as well as dense and completely grounded bounding box annotations. To achieve more explainable answers, we label the objects appeared in the complex QA pairs with different semantic types, including answer-grounded objects (both appeared and not appeared in the questions), and contextual objects for answer-grounded objects. We also propose a new 3D VQA framework to effectively predict the completely visually grounded and explainable answer. Extensive experiments verify that our newly collected benchmark datasets can be effectively used to evaluate various 3D VQA methods from different aspects and our newly proposed framework also achieves the state-of-the-art performance on the new benchmark dataset. The datasets and the source code are available via <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://github.com/zlccccc/3DVL_Codebase</uri> .
32
PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation
Raffaelli Charles, Hao Su, Kaichun Mo et al. · 2017 · 9.6K citations
Exploring the Limits of Transfer Learning with a Unified Text-to-Text\n Transformer
Colin Raffel, Noam Shazeer, Adam Roberts et al. · arXiv (Cornell University) · 2019 · 8.3K citations · Full text
FCOS: Fully Convolutional One-Stage Object Detection
Zhi Tian, Chunhua Shen, Hao Chen et al. · 2019 · 5.9K citations
Convolutional Neural Network, Image Analysis, Machine Vision +13
VQA: Visual Question Answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu et al. · 2015 · 4.2K citations
Artificial Intelligence, Natural Language Processing, Multimodal Llm +15