paper-with-me

3D Object Captioning

1개 벤치마크 · 논문 8편 · 이 태스크의 논문 보기 →

Benchmarks

Objaverse

결과 6개

Most implemented

Papers

PointAlign: Feature-Level Alignment Regularization for 3D Vision-Language Models

2026-02-28 · Yuanhao Su, Shaofeng Zhang, Xiaosong Jia, Qi Fan arxiv

The development of 3D Vision-Language Models (VLMs), crucial for applications in robotics, autonomous driving, and augmented reality, is severely constrained by the scarcity of paired 3D-text data. Existing methods rely …

3D Object CaptioningAutonomous Driving

PiSA: A Self-Augmented Data Engine and Training Strategy for 3D Understanding with Large Models

2025-03-13 · Zilu Guo, Hongbin Lin, Zhihao Yuan, Chaoda Zheng 외

3D Multimodal Large Language Models (MLLMs) have recently made substantial advancements. However, their potential remains untapped, primarily due to the limited quantity and suboptimal quality of 3D datasets. Current app…

3D Object Captioning

MARVEL-40M+: Multi-Level Visual Elaboration for High-Fidelity Text-to-3D Content Creation

2024-11-26 · CVPR 2025 1 · Sankalp Sinha, Mohammad Sadil Khan, Muhammad Usama, Shino Sam 외

Generating high-fidelity 3D content from text prompts remains a significant challenge in computer vision due to the limited size, diversity, and annotation depth of the existing datasets. To address this, we introduce MA…

3D dense captioning3D Object Captioning3D ReconstructionDiversity+2

MiniGPT-3D: Efficiently Aligning 3D Point Clouds with Large Language Models using 2D Priors

2024-05-02 · Yuan Tang, Xu Han, Xianzhi Li, Qiao Yu 외

Large 2D vision-language models (2D-LLMs) have gained significant attention by bridging Large Language Models (LLMs) with images using a simple projector. Inspired by their success, large 3D point cloud-language models (…

3D Object Captioning3D Object ClassificationGenerative 3D Object ClassificationGPU+1

View Selection for 3D Captioning via Diffusion Ranking

2024-04-11 · Tiange Luo, Justin Johnson, Honglak Lee

Scalable annotation approaches are crucial for constructing extensive 3D-text datasets, facilitating a broader range of applications. However, existing methods sometimes lead to the generation of hallucinated captions, c…

3D Object CaptioningHallucinationImage CaptioningQuestion Answering+2

ShapeLLM: Universal 3D Object Understanding for Embodied Interaction

2024-02-27 · Zekun Qi, Runpei Dong, Shaochen Zhang, Haoran Geng 외

This paper presents ShapeLLM, the first 3D Multimodal Large Language Model (LLM) designed for embodied interaction, exploring a universal 3D object understanding with 3D point clouds and languages. ShapeLLM is built upon…

3D geometry3D Object Captioning3D Point Cloud Classification3D Point Cloud Linear Classification+13

전체 8편 보기 →