paper-with-me

Papers

Visuospatial Cognitive Assistant

2025-05-18 · Qi Feng

Video-based spatial cognition is vital for robotics and embodied AI but challenges current Vision-Language Models (VLMs). This paper makes two key contributions. First, we introduce ViCA (Visuospatial Cognitive Assistant)-322K, a diverse dataset of 322,003 QA pairs from real-world indoor videos (ARKitScenes, ScanNet, ScanNet++), offering supervision for 3D metadata-grounded queries and video-based complex reasoning. Second, we develop ViCA-7B, fine-tuned on ViCA-322K, which achieves new state-of-the-art on all eight VSI-Bench tasks, outperforming existing models, including larger ones (e.g., +26.1 on Absolute Distance). For interpretability, we present ViCA-Thinking-2.68K, a dataset with explicit reasoning chains, and fine-tune ViCA-7B to create ViCA-7B-Thinking, a model that articulates its spatial reasoning. Our work highlights the importance of targeted data and suggests paths for improved temporal-spatial modeling. We release all resources to foster research in robust visuospatial intelligence.

📄 PDF Abstract BibTeX arXiv:2505.12312

Code (1)

nkkbr/vica pytorch

Tasks

Spatial Reasoning

Similar Papers 제목 키워드 기반

Towards Visuospatial Cognition via Hierarchical Fusion of Visual Experts

2025-05-18 · Qi Feng

While Multimodal Large Language Models (MLLMs) excel at general vision-language tasks, visuospatial cognition - reasoning about spatial layouts, relations, and dynamics - remains a significant challenge. Existing models …

Spatial Reasoning

Towards a Human-Centred Cognitive Model of Visuospatial Complexity in Everyday Driving

2020-05-29 · Vasiliki Kondyli, Mehul Bhatt, Jakob Suchan

We develop a human-centred, cognitive model of visuospatial complexity in everyday, naturalistic driving conditions. With a focus on visual perception, the model incorporates quantitative, structural, and dynamic attribu…

Benchmarking

Mind's Eye: A Benchmark of Visual Abstraction, Transformation and Composition for Multimodal LLMs

2026-04-17 · Rohit Sinha, Aditya Kanade, Sai Srinivas Kancheti, Vineeth N Balasubramanian 외 arxiv

Multimodal large language models (MLLMs) have achieved impressive progress on vision language benchmarks, yet their capacity for visual cognitive and visuospatial reasoning remains less understood. We introduce "Mind's E…

A Visuospatial Dataset for Naturalistic Verb Learning

2020-10-28 · Joint Conference on Lexical and Computational Semantics 2020 · Dylan Ebert, Ellie Pavlick

We introduce a new dataset for training and evaluating grounded language models. Our data is collected within a virtual reality environment and is designed to emulate the quality of language data to which a pre-verbal ch…

Grounded language learningLanguage Acquisition

Expanding the phenotype of SCA19/22: Parkinsonism, cognitive impairment and epilepsy

2020-11-20 · Vincent Huin, Isabelle Strubi-Vuillaume, Kathy Dujardin, Marine Brion 외

BACKGROUND: Spinocerebellar ataxia types 19 and 22 (SCA19/22) are rare conditions in which relatively isolated cerebellar involvement is frequently associated with cognitive impairment. Here, we report on new clinical fe…