3D Question Answering (3D-QA)
3개 벤치마크 · 논문 22편 · 이 태스크의 논문 보기 →
Benchmarks
Most implemented
Visual Instruction Tuning
Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following
3D-LLM: Injecting the 3D World into Large Language Models
ShapeLLM: Universal 3D Object Understanding for Embodied Interaction
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
PointLLM: Empowering Large Language Models to Understand Point Clouds
Papers
Unveiling the Mist over 3D Vision-Language Understanding: Object-centric Evaluation with Chain-of-Analysis
Existing 3D vision-language (3D-VL) benchmarks fall short in evaluating 3D-VL models, creating a "mist" that obscures rigorous insights into model capabilities and 3D-VL tasks. This mist persists due to three key limitat…
3D Question Answering (3D-QA)3D visual groundingDSPNet: Dual-vision Scene Perception for Robust 3D Question Answering
3D Question Answering (3D QA) requires the model to comprehensively understand its situated 3D scene described by the text, then reason about its surrounding environment and answer a question under that situation. Howeve…
3D Question Answering (3D-QA)Question AnsweringVideo-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
The rapid advancement of Multimodal Large Language Models (MLLMs) has significantly impacted various multimodal tasks. However, these models face challenges in tasks that require spatial understanding within 3D environme…
3D Question Answering (3D-QA)PositionScene UnderstandingVideo Instruction Tuning With Synthetic Data
The development of video large multimodal models (LMMs) has been hindered by the difficulty of curating large amounts of high-quality raw data from the web. To address this, we propose an alternative approach by creating…
3D Question Answering (3D-QA)Instruction FollowingMultiple-choice+5LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness
Recent advancements in Large Multimodal Models (LMMs) have greatly enhanced their proficiency in 2D visual understanding tasks, enabling them to effectively process and understand images and videos. However, the developm…
3D Question Answering (3D-QA)PositionScene UnderstandingMulti-modal Situated Reasoning in 3D Scenes
Situation awareness is essential for understanding and reasoning about 3D scenes in embodied AI agents. However, existing datasets and benchmarks for situated understanding are limited in data modality, diversity, scale,…
3D Question Answering (3D-QA)