paper-with-me

홈 › Papers

Beyond Human Perception: Understanding Multi-Object World from Monocular View

2025-01-01 · CVPR 2025 1 · Keyu Guo, Yongle Huang, ShiJie Sun, XiangYu Song, Mingtao Feng, Zedong Liu, HuanSheng Song, Tiantian Wang, JianXin Li, Naveed Akhtar, Ajmal Saeed Mian

Language and binocular vision play a crucial role in human understanding of the world. Advancements in artificial intelligence have also made it possible for machines to develop 3D perception capabilities essential for high-level scene understanding. However, only monocular cameras are often available in practice due to cost and space constraints. Enabling machines to achieve accurate 3D understanding from a monocular view is practical but presents significant challenges. We introduce MonoMulti-3DVG, a novel task aimed at achieving multi-object 3D Visual Grounding (3DVG) based on monocular RGB images, allowing machines to better understand and interact with the 3D world. To this end, we construct a large-scale benchmark dataset, MonoMulti3D-ROPE, and propose a model, CyclopsNet that integrates a State-Prompt Visual Encoder (SPVE) module with a Denoising Alignment Fusion (DAF) module to achieve robust multi-modal semantic alignment and fusion. This leads to more stable and robust multi-modal joint representations for downstream tasks. Experimental results show that our method significantly outperforms existing techniques on the MonoMulti3D-ROPE dataset. Our dataset and code are available at https://github.com/JasonHuang516/MonoMulti-3DVG

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

3D visual groundingDenoisingScene UnderstandingVisual Grounding

Similar Papers 제목 키워드 기반

Human-Centric Foundation Models: Perception, Generation and Agentic Modeling

2025-02-12 · Shixiang Tang, Yizhou Wang, Lu Chen, YuAn Wang 외

Human understanding and generation are critical for modeling digital humans and humanoid embodiments. Recently, Human-centric Foundation Models (HcFMs) inspired by the success of generalist models, such as large language…

Survey

Visual Grounding from Event Cameras

2025-09-11 · Lingdong Kong, Dongyue Lu, Ao Liang, Rong Li 외 arxiv

Event cameras capture changes in brightness with microsecond precision and remain reliable under motion blur and challenging illumination, offering clear advantages for modeling highly dynamic scenes. Yet, their integrat…

Natural Language UnderstandingObject RecognitionVisual Grounding

Do large language vision models understand 3D shapes?

2024-12-14 · Sagi Eppel

Large vision language models (LVLM) are the leading A.I approach for achieving a general visual understanding of the world. Models such as GPT, Claude, Gemini, and LLama can use images to understand and analyze complex v…

Beyond Planar Symmetry: Modeling human perception of reflection and rotation symmetries in the wild

2017-04-11 · ICCV 2017 10 · Christopher Funk, Yanxi Liu

Humans take advantage of real world symmetries for various tasks, yet capturing their superb symmetry perception mechanism with a computational model remains elusive. Motivated by a new study demonstrating the extremely …

Symmetry Detection

Looking Beyond the Visible Scene

2014-06-01 · CVPR 2014 6 · Aditya Khosla, Byoungkwon An An, Joseph J. Lim, Antonio Torralba

A common thread that ties together many prior works in scene understanding is their focus on the aspects directly present in a scene such as its categorical classification or the set of objects. In this work, we propose …

Scene Understanding