paper-with-me

홈 › Papers

From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D

2025-03-29 · Jiahui Zhang, Yurui Chen, Yanpeng Zhou, Yueming Xu, Ze Huang, Jilin Mei, Junhui Chen, Yu-Jie Yuan, Xinyue Cai, Guowei Huang, Xingyue Quan, Hang Xu, Li Zhang

Recent advances in LVLMs have improved vision-language understanding, but they still struggle with spatial perception, limiting their ability to reason about complex 3D scenes. Unlike previous approaches that incorporate 3D representations into models to improve spatial understanding, we aim to unlock the potential of VLMs by leveraging spatially relevant image data. To this end, we introduce a novel 2D spatial data generation and annotation pipeline built upon scene data with 3D ground-truth. This pipeline enables the creation of a diverse set of spatial tasks, ranging from basic perception tasks to more complex reasoning tasks. Leveraging this pipeline, we construct SPAR-7M, a large-scale dataset generated from thousands of scenes across multiple public datasets. In addition, we introduce SPAR-Bench, a benchmark designed to offer a more comprehensive evaluation of spatial capabilities compared to existing spatial benchmarks, supporting both single-view and multi-view inputs. Training on both SPAR-7M and large-scale 2D datasets enables our models to achieve state-of-the-art performance on 2D spatial benchmarks. Further fine-tuning on 3D task-specific datasets yields competitive results, underscoring the effectiveness of our dataset in enhancing spatial reasoning.

📄 PDF Abstract BibTeX arXiv:2503.22976

Code (1)

fudan-zvg/spar pytorch

Tasks

Spatial Reasoning

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

FlatLand: Personalized Graph Federated Learning via Tailored Lorentz Space

2026-08-21 · Jiahong Liu, Ram Samarth B B, Xinyu Fu, Menglin Yang 외 arxiv

Federated learning enables privacy-preserving collaborative training, but highly heterogeneous client data remain challenging, especially in graph federated learning where clients possess structurally diverse graphs. Exi…

Personalized Federated LearningGraph Learning

Flatland: a Lightweight First-Person 2-D Environment for Reinforcement Learning

2018-09-03 · Hugo Caselles-Dupré, Louis Annabi, Oksana Hagen, Michael Garcia-Ortiz 외

Flatland is a simple, lightweight environment for fast prototyping and testing of reinforcement learning agents. It is of lower complexity compared to similar 3D platforms (e.g. DeepMind Lab or VizDoom), but emulates phy…

Lifelong learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Flatland-RL : Multi-Agent Reinforcement Learning on Trains

2020-12-10 · Sharada Mohanty, Erik Nygren, Florian Laurent, Manuel Schneider 외

Efficient automated scheduling of trains remains a major challenge for modern railway systems. The underlying vehicle rescheduling problem (VRSP) has been a major focus of Operations Research (OR) since decades. Traditio…

Imitation LearningMulti-agent Reinforcement Learningreinforcement-learningReinforcement Learning+2

Scope Restriction for Scalable Real-Time Railway Rescheduling: An Exploratory Study

2023-05-05 · Erik Nygren, Christian Eichenberger, Emma Frejinger

With the aim to stimulate future research, we describe an exploratory study of a railway rescheduling problem. A widely used approach in practice and state of the art is to decompose these complex problems by geographica…

3DoF Localization from a Single Image and an Object Map: the Flatlandia Problem and Dataset

2023-04-13 · Matteo Toso, Matteo Taiana, Stuart James, Alessio Del Bue

Efficient visual localization is crucial to many applications, such as large-scale deployment of autonomous agents and augmented reality. Traditional visual localization, while achieving remarkable accuracy, relies on ex…

Privacy PreservingVisual Localization