3D dense captioning
2개 벤치마크 · 논문 33편 · 이 태스크의 논문 보기 →
Benchmarks
ScanRefer Dataset
Nr3D
Most implemented
MARVEL-40M+: Multi-Level Visual Elaboration for High-Fidelity Text-to-3D Content Creation
3D CoCa: Contrastive Learners are 3D Captioners
Rethinking 3D Dense Caption and Visual Grounding in A Unified Framework through Prompt-based Localization
TOD3Cap: Towards 3D Dense Captioning in Outdoor Scenes
Papers
3D-Aware VLMs with Implicit and Explicit Geometries
Despite rapid progress, most existing vision-language models (VLMs) built from 2D visual inputs often struggle when handling various 3D tasks that require fine-grained spatial understanding and reasoning. To bridge this …
3D dense captioningSpatial ReasoningVisual GroundingPVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet
3D dense captioning, an emerging vision-language task, aims to generate descriptive sentences for each object in the 3D scene. Despite the impressive results achieved by previous methods, they suffer from two limitations…
3D dense captioningData AugmentationOcc-VLM: Occupancy Grounded Vision Language Model for Indoor Scene Understanding
Recently, vision-language models (VLMs) have made significant progress in 3D scene understanding, driving advances in applications such as embodied intelligence and robotic vision. However, existing approaches typically …
Visual Question Answering3D dense captioningScene UnderstandingPoint CloudsGAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models
Multimodal Large Language Models (MLLMs) demonstrate exceptional semantic reasoning but struggle with 3D spatial perception when restricted to pure RGB inputs. Despite leveraging implicit geometric priors from 3D reconst…
Video Object Detection3D dense captioning3D ReconstructionVisual GroundingScenes as Tokens: Multi-Scale Normal Distributions Transform Tokenizer for General 3D Vision-Language Understanding
Recent advances in 3D vision-language models (VLMs) highlight a strong potential for 3D scene understanding and reasoning. However, effectively tokenizing 3D scenes into holistic scene tokens, and leveraging these tokens…
Visual Question Answering3D dense captioningScene UnderstandingPoint CloudsVid-LLM: A Compact Video-based 3D Multimodal LLM with Reconstruction-Reasoning Synergy
Recent developments in Multimodal Large Language Models (MLLMs) have significantly improved Vision-Language (VL) reasoning in 2D domains. However, extending these capabilities to 3D scene understanding remains a major ch…
3D dense captioningScene UnderstandingQuestion AnsweringVisual Grounding