Generative 3D Object Classification
2개 벤치마크 · 논문 5편 · 이 태스크의 논문 보기 →
Benchmarks
Objaverse
ModelNet40
Most implemented
Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following
3D-LLM: Injecting the 3D World into Large Language Models
ShapeLLM: Universal 3D Object Understanding for Embodied Interaction
PointLLM: Empowering Large Language Models to Understand Point Clouds
Papers
MiniGPT-3D: Efficiently Aligning 3D Point Clouds with Large Language Models using 2D Priors
Large 2D vision-language models (2D-LLMs) have gained significant attention by bridging Large Language Models (LLMs) with images using a simple projector. Inspired by their success, large 3D point cloud-language models (…
3D Object Captioning3D Object ClassificationGenerative 3D Object ClassificationGPU+1ShapeLLM: Universal 3D Object Understanding for Embodied Interaction
This paper presents ShapeLLM, the first 3D Multimodal Large Language Model (LLM) designed for embodied interaction, exploring a universal 3D object understanding with 3D point clouds and languages. ShapeLLM is built upon…
3D geometry3D Object Captioning3D Point Cloud Classification3D Point Cloud Linear Classification+13Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following
We introduce Point-Bind, a 3D multi-modality model aligning point clouds with 2D image, language, audio, and video. Guided by ImageBind, we construct a joint embedding space between 3D and multi-modalities, enabling many…
3D Generation3D Question Answering (3D-QA)Generative 3D Object ClassificationInstruction Following+5PointLLM: Empowering Large Language Models to Understand Point Clouds
The unprecedented advancements in Large Language Models (LLMs) have shown a profound impact on natural language processing but are yet to fully embrace the realm of 3D understanding. This paper introduces PointLLM, a pre…
3D Object Captioning3D Object Classification3D Question Answering (3D-QA)Common Sense Reasoning+23D-LLM: Injecting the 3D World into Large Language Models
Large language models (LLMs) and Vision-Language Models (VLMs) have been proven to excel at multiple tasks, such as commonsense reasoning. Powerful as these models can be, they are not grounded in the 3D physical world, …
3D Object Captioning3D Question Answering (3D-QA)Dense CaptioningGenerative 3D Object Classification+1