paper-with-me

홈 › Papers

Comprehensive Visual Question Answering on Point Clouds through Compositional Scene Manipulation

2021-12-22 · Xu Yan, Zhihao Yuan, Yuhao Du, Yinghong Liao, Yao Guo, Zhen Li, Shuguang Cui

Visual Question Answering on 3D Point Cloud (VQA-3D) is an emerging yet challenging field that aims at answering various types of textual questions given an entire point cloud scene. To tackle this problem, we propose the CLEVR3D, a large-scale VQA-3D dataset consisting of 171K questions from 8,771 3D scenes. Specifically, we develop a question engine leveraging 3D scene graph structures to generate diverse reasoning questions, covering the questions of objects' attributes (i.e., size, color, and material) and their spatial relationships. Through such a manner, we initially generated 44K questions from 1,333 real-world scenes. Moreover, a more challenging setup is proposed to remove the confounding bias and adjust the context from a common-sense layout. Such a setup requires the network to achieve comprehensive visual understanding when the 3D scene is different from the general co-occurrence context (e.g., chairs always exist with tables). To this end, we further introduce the compositional scene manipulation strategy and generate 127K questions from 7,438 augmented 3D scenes, which can improve VQA-3D models for real-world comprehension. Built upon the proposed dataset, we baseline several VQA-3D models, where experimental results verify that the CLEVR3D can significantly boost other 3D scene understanding tasks. Our code and dataset will be made publicly available at https://github.com/yanx27/CLEVR3D.

📄 PDF Abstract BibTeX arXiv:2112.11691

Code (1)

yanx27/clevr3d 공식 구현 pytorch

Tasks

Common Sense ReasoningQuestion AnsweringScene UnderstandingVisual Question AnsweringVisual Question Answering (VQA)Visual Reasoning

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Position-Wise Feed-Forward Layer 설명 없음
Adam 설명 없음
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…

Similar Papers 제목 키워드 기반

DSPNet: Dual-vision Scene Perception for Robust 3D Question Answering

2025-03-05 · CVPR 2025 1 · Jingzhou Luo, Yang Liu, Weixing Chen, Zhen Li 외

3D Question Answering (3D QA) requires the model to comprehensively understand its situated 3D scene described by the text, then reason about its surrounding environment and answer a question under that situation. Howeve…

3D Question Answering (3D-QA)Question Answering

NuScenes-QA: A Multi-modal Visual Question Answering Benchmark for Autonomous Driving Scenario

2023-05-24 · Tianwen Qian, Jingjing Chen, Linhai Zhuo, Yang Jiao 외

We introduce a novel visual question answering (VQA) task in the context of autonomous driving, aiming to answer natural language questions based on street-view clues. Compared to traditional VQA tasks, VQA in autonomous…

Autonomous DrivingQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Look Around and Refer: 2D Synthetic Semantics Knowledge Distillation for 3D Visual Grounding

2022-11-25 · Eslam Mohamed BAKR, Yasmeen Alsaedy, Mohamed Elhoseiny

The 3D visual grounding task has been explored with visual and language streams comprehending referential language to identify target objects in 3D scenes. However, most existing methods devote the visual stream to captu…

3D visual groundingKnowledge DistillationVisual Grounding

Embodied Question Answering in Photorealistic Environments with Point Cloud Perception

2019-04-06 · CVPR 2019 6 · Erik Wijmans, Samyak Datta, Oleksandr Maksymets, Abhishek Das 외

To help bridge the gap between internet vision-style problems and the goal of vision for embodied perception we instantiate a large-scale navigation task -- Embodied Question Answering [1] in photo-realistic environments…

Embodied Question AnsweringQuestion Answering

Recent, rapid advancement in visual question answering architecture: a review

2022-03-02 · Venkat Kodali, Daniel Berleant

Understanding visual question answering is going to be crucial for numerous human activities. However, it presents major challenges at the heart of the artificial intelligence endeavor. This paper presents an update on t…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)