3D Object Captioning
1개 벤치마크 · 논문 8편 · 이 태스크의 논문 보기 →
Benchmarks
Objaverse
Most implemented
3D-LLM: Injecting the 3D World into Large Language Models
ShapeLLM: Universal 3D Object Understanding for Embodied Interaction
PointLLM: Empowering Large Language Models to Understand Point Clouds
MARVEL-40M+: Multi-Level Visual Elaboration for High-Fidelity Text-to-3D Content Creation
View Selection for 3D Captioning via Diffusion Ranking
Papers
PointAlign: Feature-Level Alignment Regularization for 3D Vision-Language Models
The development of 3D Vision-Language Models (VLMs), crucial for applications in robotics, autonomous driving, and augmented reality, is severely constrained by the scarcity of paired 3D-text data. Existing methods rely …
3D Object CaptioningAutonomous DrivingPiSA: A Self-Augmented Data Engine and Training Strategy for 3D Understanding with Large Models
3D Multimodal Large Language Models (MLLMs) have recently made substantial advancements. However, their potential remains untapped, primarily due to the limited quantity and suboptimal quality of 3D datasets. Current app…
3D Object CaptioningMARVEL-40M+: Multi-Level Visual Elaboration for High-Fidelity Text-to-3D Content Creation
Generating high-fidelity 3D content from text prompts remains a significant challenge in computer vision due to the limited size, diversity, and annotation depth of the existing datasets. To address this, we introduce MA…
3D dense captioning3D Object Captioning3D ReconstructionDiversity+2MiniGPT-3D: Efficiently Aligning 3D Point Clouds with Large Language Models using 2D Priors
Large 2D vision-language models (2D-LLMs) have gained significant attention by bridging Large Language Models (LLMs) with images using a simple projector. Inspired by their success, large 3D point cloud-language models (…
3D Object Captioning3D Object ClassificationGenerative 3D Object ClassificationGPU+1View Selection for 3D Captioning via Diffusion Ranking
Scalable annotation approaches are crucial for constructing extensive 3D-text datasets, facilitating a broader range of applications. However, existing methods sometimes lead to the generation of hallucinated captions, c…
3D Object CaptioningHallucinationImage CaptioningQuestion Answering+2ShapeLLM: Universal 3D Object Understanding for Embodied Interaction
This paper presents ShapeLLM, the first 3D Multimodal Large Language Model (LLM) designed for embodied interaction, exploring a universal 3D object understanding with 3D point clouds and languages. ShapeLLM is built upon…
3D geometry3D Object Captioning3D Point Cloud Classification3D Point Cloud Linear Classification+13