paper-with-me

홈 › Papers

3D-GRAND: A Million-Scale Dataset for 3D-LLMs with Better Grounding and Less Hallucination

2024-06-07 · CVPR 2025 1 · Jianing Yang, Xuweiyi Chen, Nikhil Madaan, Madhavan Iyengar, Shengyi Qian, David F. Fouhey, Joyce Chai

The integration of language and 3D perception is crucial for embodied agents and robots that comprehend and interact with the physical world. While large language models (LLMs) have demonstrated impressive language understanding and generation capabilities, their adaptation to 3D environments (3D-LLMs) remains in its early stages. A primary challenge is a lack of large-scale datasets with dense grounding between language and 3D scenes. We introduce 3D-GRAND, a pioneering large-scale dataset comprising 40,087 household scenes paired with 6.2 million densely-grounded scene-language instructions. Our results show that instruction tuning with 3D-GRAND significantly enhances grounding capabilities and reduces hallucinations in 3D-LLMs. As part of our contributions, we propose a comprehensive benchmark 3D-POPE to systematically evaluate hallucination in 3D-LLMs, enabling fair comparisons of models. Our experiments highlight a scaling effect between dataset size and 3D-LLM performance, emphasizing the importance of large-scale 3D-text datasets for embodied AI research. Our results demonstrate early signals for effective sim-to-real transfer, indicating that models trained on large synthetic data can perform well on real-world 3D scans. Through 3D-GRAND and 3D-POPE, we aim to equip the embodied AI community with resources and insights to lead to more reliable and better-grounded 3D-LLMs. Project website: https://3d-grand.github.io

📄 PDF Abstract BibTeX arXiv:2406.05132

Code (1)

sled-group/3D-GRAND 공식 구현 pytorch

Tasks

Hallucination

Similar Papers 제목 키워드 기반

GRAND+: Scalable Graph Random Neural Networks

2022-03-12 · Wenzheng Feng, Yuxiao Dong, Tinglin Huang, Ziqi Yin 외

Graph neural networks (GNNs) have been widely adopted for semi-supervised learning on graphs. A recent study shows that the graph random neural network (GRAND) model can generate state-of-the-art performance for this pro…

Data AugmentationGraph LearningModel OptimizationNode Classification

Amortized Planning with Large-Scale Transformers: A Case Study on Chess

2024-02-07 · Anian Ruoss, Grégoire Delétang, Sourabh Medapati, Jordi Grau-Moya 외

This paper uses chess, a landmark planning problem in AI, to assess transformers' performance on a planning task where memorization is futile $\unicode{x2013}$ even at a large scale. To this end, we release ChessBench, a…

Memorization

ANALOGYKB: Unlocking Analogical Reasoning of Language Models with A Million-scale Knowledge Base

2023-05-10 · Siyu Yuan, Jiangjie Chen, Changzhi Sun, Jiaqing Liang 외

Analogical reasoning is a fundamental cognitive ability of humans. However, current language models (LMs) still struggle to achieve human-like performance in analogical reasoning tasks due to a lack of resources for mode…

Knowledge Graphs

WinoWhat: A Parallel Corpus of Paraphrased WinoGrande Sentences with Common Sense Categorization

2025-03-31 · Ine Gevers, Victor De Marez, Luna De Bruyne, Walter Daelemans

In this study, we take a closer look at how Winograd schema challenges can be used to evaluate common sense reasoning in LLMs. Specifically, we evaluate generative models of different sizes on the popular WinoGrande benc…

Common Sense ReasoningMemorizationWinogrande

Graph Random Neural Networks for Semi-Supervised Learning on Graphs

2020-12-01 · NeurIPS 2020 12 · Wenzheng Feng, Jie Zhang, Yuxiao Dong, Yu Han 외

We study the problem of semi-supervised learning on graphs, for which graph neural networks (GNNs) have been extensively explored. However, most existing GNNs inherently suffer from the limitations of over-smoothing, non…

Data AugmentationNode Classification