CVP: Central-Peripheral Vision-Inspired Multimodal Model for Spatial Reasoning
We present a central-peripheral vision-inspired framework (CVP), a simple yet effective multimodal model for spatial reasoning that draws inspiration from the two types of human visual fields -- central vision and peripheral vision. Existing approaches primarily rely on unstructured representations, such as point clouds, voxels, or patch features, and inject scene context implicitly via coordinate embeddings. However, this often results in limited spatial reasoning capabilities due to the lack of explicit, high-level structural understanding. To address this limitation, we introduce two complementary components into a Large Multimodal Model-based architecture: target-affinity token, analogous to central vision, that guides the model's attention toward query-relevant objects; and allocentric grid, akin to peripheral vision, that captures global scene context and spatial arrangements. These components work in tandem to enable structured, context-aware understanding of complex 3D environments. Experiments show that CVP achieves state-of-the-art performance across a range of 3D scene understanding benchmarks.
Code (0)
등록된 구현이 없습니다.
Tasks
Scene UnderstandingSpatial ReasoningPoint CloudsSimilar Papers 제목 키워드 기반
How spatial frequencies and color drive object search in real-world scenes: A new eye-movement corpus
When studying how people search for objects in scenes, the inhomogeneity of the visual field is often ignored. Due to physiological limitations peripheral vision is blurred and mainly uses coarse-grained information (i.e…
ObjectObject LocalizationSpatial frequency processing in the central and peripheral visual field during scene viewing
Visuospatial attention and gaze control depend on the interaction of foveal and peripheral processing. The foveal and peripheral regions of the visual field are differentially sensitive to parts of the spatial-frequency …
Modeling the Contribution of Central Versus Peripheral Vision in Scene, Object, and Face Recognition
It is commonly believed that the central visual field is important for recognizing objects and faces, and the peripheral region is useful for scene recognition. However, the relative importance of central versus peripher…
Face RecognitionObject RecognitionScene RecognitionA Gated Peripheral-Foveal Convolutional Neural Network for Unified Image Aesthetic Prediction
Learning fine-grained details is a key issue in image aesthetic assessment. Most of the previous methods extract the fine-grained details via random cropping strategy, which may undermine the integrity of semantic inform…
AMuSE: Adaptive Multimodal Analysis for Speaker Emotion Recognition in Group Conversations
Analyzing individual emotions during group conversation is crucial in developing intelligent agents capable of natural human-machine interaction. While reliable emotion recognition techniques depend on different modaliti…
Emotion Recognition