paper-with-me

Papers

3DCoMPaT200: Language-Grounded Compositional Understanding of Parts and Materials of 3D Shapes

2025-01-12 · Mahmoud Ahmed, Xiang Li, Arpit Prajapati, Mohamed Elhoseiny

Understanding objects in 3D at the part level is essential for humans and robots to navigate and interact with the environment. Current datasets for part-level 3D object understanding encompass a limited range of categories. For instance, the ShapeNet-Part and PartNet datasets only include 16, and 24 object categories respectively. The 3DCoMPaT dataset, specifically designed for compositional understanding of parts and materials, contains only 42 object categories. To foster richer and fine-grained part-level 3D understanding, we introduce 3DCoMPaT200, a large-scale dataset tailored for compositional understanding of object parts and materials, with 200 object categories with $\approx$5 times larger object vocabulary compared to 3DCoMPaT and $\approx$ 4 times larger part categories. Concretely, 3DCoMPaT200 significantly expands upon 3DCoMPaT, featuring 1,031 fine-grained part categories and 293 distinct material classes for compositional application to 3D object parts. Additionally, to address the complexities of compositional 3D modeling, we propose a novel task of Compositional Part Shape Retrieval using ULIP to provide a strong 3D foundational model for 3D Compositional Understanding. This method evaluates the model shape retrieval performance given one, three, or six parts described in text format. These results show that the model's performance improves with an increasing number of style compositions, highlighting the critical role of the compositional dataset. Such results underscore the dataset's effectiveness in enhancing models' capability to understand complex 3D shapes from a compositional perspective. Code and Data can be found at http://github.com/3DCoMPaT200/3DCoMPaT200

📄 PDF Abstract BibTeX arXiv:2501.06785

Code (1)

3dcompat200/3dcompat200 공식 구현 pytorch

Tasks

NavigateObjectRetrieval

Similar Papers 제목 키워드 기반

3DCoMPaT$^{++}$: An improved Large-scale 3D Vision Dataset for Compositional Recognition

2023-10-27 · Habib Slim, Xiang Li, Yuchen Li, Mahmoud Ahmed 외

In this work, we present 3DCoMPaT$^{++}$, a multimodal 2D/3D dataset with 160 million rendered views of more than 10 million stylized 3D shapes carefully annotated at the part-instance level, alongside matching RGB point…

Kestrel: Point Grounding Multimodal LLM for Part-Aware 3D Vision-Language Understanding

2024-05-29 · Junjie Fei, Mahmoud Ahmed, Jian Ding, Eslam Mohamed BAKR 외

While 3D MLLMs have achieved significant progress, they are restricted to object and scene understanding and struggle to understand 3D spatial structures at the part level. In this paper, we introduce Kestrel, representi…

Scene UnderstandingSegmentation

A Benchmark for Systematic Generalization in Grounded Language Understanding

2020-03-11 · NeurIPS 2020 12 · Laura Ruis, Jacob Andreas, Marco Baroni, Diane Bouchacourt 외

Humans easily interpret expressions that describe unfamiliar situations composed from familiar parts ("greet the pink brontosaurus by the ferris wheel"). Modern neural networks, by contrast, struggle to interpret novel c…

Systematic Generalization

3D-PLOT-LLM: Part-Level Object Tokens for 3D Large Language Models

2026-06-18 · Jintang Xue, Xinyu Wang, Yixing Wu, Jingwen Chen 외 arxiv

3D multimodal large language models (3D MLLMs) describe a 3D object as a whole but cannot address, name, or reason about its parts. Prior part-aware attempts add segmentation decoders, heavier 3D encoders, or bounding-bo…

Compositional Networks Enable Systematic Generalization for Grounded Language Understanding

2020-08-06 · Findings (EMNLP) 2021 11 · Yen-Ling Kuo, Boris Katz, Andrei Barbu

Humans are remarkably flexible when understanding new sentences that include combinations of concepts they have never encountered before. Recent work has shown that while deep networks can mimic some human language abili…

Systematic Generalization