3D Question Answering

被引:2
|
作者
Ye, Shuquan [1 ]
Chen, Dongdong [2 ]
Han, Songfang [3 ]
Liao, Jing [1 ]
机构
[1] City Univ Hong Kong, Kowloon Tong, Hong Kong, Peoples R China
[2] Microsoft Cloud AI, Redmond, WA 98052 USA
[3] Univ Calif San Diego, La Jolla, CA 92093 USA
关键词
Point cloud; scene understanding; LANGUAGE; VISION;
D O I
10.1109/TVCG.2022.3225327
中图分类号
TP31 [计算机软件];
学科分类号
081202 ; 0835 ;
摘要
Visual question answering (VQA) has experienced tremendous progress in recent years. However, most efforts have only focused on 2D image question-answering tasks. In this article, we extend VQA to its 3D counterpart, 3D question answering (3DQA), which can facilitate a machine's perception of 3D real-world scenarios. Unlike 2D image VQA, 3DQA takes the color point cloud as input and requires both appearance and 3D geometrical comprehension to answer the 3D-related questions. To this end, we propose a novel transformer-based 3DQA framework "3DQA-TR", which consists of two encoders to exploit the appearance and geometry information, respectively. Finally, the multi-modal information about the appearance, geometry, and linguistic question can attend to each other via a 3D-linguistic Bert to predict the target answers. To verify the effectiveness of our proposed 3DQA framework, we further develop the first 3DQA dataset "ScanQA", which builds on the ScanNet dataset and contains over 10 K question-answer pairs for 806 scenes. To the best of our knowledge, ScanQA is the first large-scale dataset with natural-language questions and free-form answers in 3D environments that is fully human-annotated. We also use several visualizations and experiments to investigate the astonishing diversity of the collected questions and the significant differences between this task from 2D VQA and 3D captioning. Extensive experiments on this dataset demonstrate the obvious superiority of our proposed 3DQA framework over state-of-the-art VQA frameworks and the effectiveness of our major designs. Our code and dataset will be made publicly available to facilitate research in this direction. The code and data are available at http://shuquanye.com/3DQA_website/.
引用
收藏
页码:1772 / 1786
页数:15
相关论文
共 50 条
  • [1] 3D Question Answering
    Ye, Shuquan
    Chen, Dongdong
    Han, Songfang
    Liao, Jing
    IEEE Transactions on Visualization and Computer Graphics, 2022, 30 (03) : 1772 - 1786
  • [2] Incorporating 3D Information into Visual Question Answering
    Qiu, Yue
    Satoh, Yutaka
    Suzuki, Ryota
    Kataoka, Hirokatsu
    2019 INTERNATIONAL CONFERENCE ON 3D VISION (3DV 2019), 2019, : 756 - 765
  • [3] 3DVQA: Visual Question Answering for 3D Environments
    Etesam, Yasaman
    Kochiev, Leon
    Chang, Angel X.
    2022 19TH CONFERENCE ON ROBOTS AND VISION (CRV 2022), 2022, : 233 - 240
  • [4] ScanQA: 3D Question Answering for Spatial Scene Understanding
    Azuma, Laichi
    Miyanishi, Taiki
    Kurita, Shuhei
    Kawanahe, Motoaki
    2022 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2022), 2022, : 19107 - 19117
  • [5] Bridging the Gap between 2D and 3D Visual Question Answering: A Fusion Approach for 3D VQA
    Mo, Wentao
    Liu, Yang
    THIRTY-EIGHTH AAAI CONFERENCE ON ARTIFICIAL INTELLIGENCE, VOL 38 NO 5, 2024, : 4261 - 4268
  • [6] Question Answering with BERT: designing a 3D virtual avatar for Cultural Heritage exploration
    Farella, Mariella
    Chiazzese, Giuseppe
    Lo Bosco, Giosue
    2022 IEEE 21ST MEDITERRANEAN ELECTROTECHNICAL CONFERENCE (IEEE MELECON 2022), 2022, : 770 - 774
  • [7] Toward Explainable 3D Grounded Visual Question Answering: A New Benchmark and Strong Baseline
    Zhao, Lichen
    Cai, Daigang
    Zhang, Jing
    Sheng, Lu
    Xu, Dong
    Zheng, Rui
    Zhao, Yinjie
    Wang, Lipeng
    Fan, Xibo
    IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS FOR VIDEO TECHNOLOGY, 2023, 33 (06) : 2935 - 2949
  • [8] Weakly-Supervised 3D Spatial Reasoning for Text-Based Visual Question Answering
    Li, Hao
    Huang, Jinfa
    Jin, Peng
    Song, Guoli
    Wu, Qi
    Chen, Jie
    IEEE TRANSACTIONS ON IMAGE PROCESSING, 2023, 32 : 3367 - 3382
  • [9] 3D-Aware Visual Question Answering about Parts, Poses and Occlusions
    Wang, Xingrui
    Ma, Wufei
    Li, Zhuowan
    Kortylewski, Adam
    Yuille, Alan
    ADVANCES IN NEURAL INFORMATION PROCESSING SYSTEMS 36, NEURIPS 2023, 2023,
  • [10] Turkish question answering - Question answering for distance education students
    Yurekli, Burcu
    Arslan, Ahmet
    Senel, Hakan G.
    Yilmazel, Ozgur
    ICSOFT 2008: PROCEEDINGS OF THE THIRD INTERNATIONAL CONFERENCE ON SOFTWARE AND DATA TECHNOLOGIES, VOL ISDM/ABF, 2008, : 320 - +