paper-with-me

3D Question Answering (3D-QA)

3개 벤치마크 · 논문 22편 · 이 태스크의 논문 보기 →

Benchmarks

ScanQA Test w/ objects

결과 18개

SQA3D

결과 13개

3D MM-Vet

결과 5개

Most implemented

Visual Instruction Tuning

2023-04-17 · 구현 13개

Papers

Unveiling the Mist over 3D Vision-Language Understanding: Object-centric Evaluation with Chain-of-Analysis

2025-03-28 · CVPR 2025 1 · Jiangyong Huang, Baoxiong Jia, Yan Wang, Ziyu Zhu 외

Existing 3D vision-language (3D-VL) benchmarks fall short in evaluating 3D-VL models, creating a "mist" that obscures rigorous insights into model capabilities and 3D-VL tasks. This mist persists due to three key limitat…

3D Question Answering (3D-QA)3D visual grounding

DSPNet: Dual-vision Scene Perception for Robust 3D Question Answering

2025-03-05 · CVPR 2025 1 · Jingzhou Luo, Yang Liu, Weixing Chen, Zhen Li 외

3D Question Answering (3D QA) requires the model to comprehensively understand its situated 3D scene described by the text, then reason about its surrounding environment and answer a question under that situation. Howeve…

3D Question Answering (3D-QA)Question Answering

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

2024-11-30 · CVPR 2025 1 · Duo Zheng, Shijia Huang, LiWei Wang

The rapid advancement of Multimodal Large Language Models (MLLMs) has significantly impacted various multimodal tasks. However, these models face challenges in tasks that require spatial understanding within 3D environme…

3D Question Answering (3D-QA)PositionScene Understanding

Video Instruction Tuning With Synthetic Data

2024-10-03 · Yuanhan Zhang, Jinming Wu, Wei Li, Bo Li 외

The development of video large multimodal models (LMMs) has been hindered by the difficulty of curating large amounts of high-quality raw data from the web. To address this, we propose an alternative approach by creating…

3D Question Answering (3D-QA)Instruction FollowingMultiple-choice+5

LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

2024-09-26 · Chenming Zhu, Tai Wang, Wenwei Zhang, Jiangmiao Pang 외

Recent advancements in Large Multimodal Models (LMMs) have greatly enhanced their proficiency in 2D visual understanding tasks, enabling them to effectively process and understand images and videos. However, the developm…

3D Question Answering (3D-QA)PositionScene Understanding

Multi-modal Situated Reasoning in 3D Scenes

2024-09-04 · Xiongkun Linghu, Jiangyong Huang, Xuesong Niu, Xiaojian Ma 외

Situation awareness is essential for understanding and reasoning about 3D scenes in embodied AI agents. However, existing datasets and benchmarks for situated understanding are limited in data modality, diversity, scale,…

3D Question Answering (3D-QA)

전체 22편 보기 →