paper-with-me

홈 › Papers

ReSiReg: Towards Spatially Consistent Semantics in Language-Conditioned Robotic Tasks

2026-06-17 · Simon Schwaiger, David Seyser, Alessandro Scherl, Wilfried Wöber, Gerald Steinbauer-Wagner arxiv

Vision-Language Models (VLMs) enable robots to follow open-language instructions. However, dense VLM embeddings have shown to be noisy and lack spatial consistency. This is problematic for robotic applications, which require simultaneous reasoning over semantics and 3D space. We examine spatial structure across recent VLMs and propose ReSiReg, a feature reconstruction method that uses spatially consistent VLM intermediates to improve dense language-grounded retrieval. ReSiReg clusters intermediates into visual prototypes, derives their language descriptors, and reconstructs each patch as a soft mixture of prototype-level language embeddings. We evaluate quantitatively on OVSS and 3D mapping across backbones, and qualitatively in real-world manipulation scenes. Quantitative results show improved dense retrieval; manipulation scenes show more spatially consistent target activations. We further provide a compact 25M dense VLM for robotic applications, substantially smaller than and competitive with ViT-B baselines. Available at https://resireg.github.io

📄 PDF Abstract BibTeX arXiv:2606.19088

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

PosA-VLA: Enhancing Action Generation via Pose-Conditioned Anchor Attention

2025-12-03 · Ziwen Li, Xin Wang, Hanlue Zhang, Runnan Chen 외 arxiv

The Vision-Language-Action (VLA) models have demonstrated remarkable performance on embodied tasks and shown promising potential for real-world applications. However, current VLAs still struggle to produce consistent and…

VectorSynth: Fine-Grained Satellite Image Synthesis with Structured Semantics

2025-11-11 · Daniel Cher, Brian Wei, Srikumar Sastry, Nathan Jacobs arxiv

We introduce VectorSynth, a diffusion-based framework for pixel-accurate satellite image synthesis conditioned on polygonal geographic annotations with semantic attributes. Unlike prior text- or layout-conditioned models…

Conditional Image Generation

DRAW2ACT: Turning Depth-Encoded Trajectories into Robotic Demonstration Videos

2025-12-16 · Yang Bai, Liudi Yang, George Eskandar, Fengyi Shen 외 arxiv

Video diffusion models provide powerful real-world simulators for embodied AI but remain limited in controllability for robotic manipulation. Recent works on trajectory-conditioned video generation address this gap but o…

Video Generation

Beyond AlphaEarth: Toward Human-Centered Geospatial Foundation Models via POI-Guided Contrastive Learning

2025-10-10 · Junyuan Liu, Quan Qin, Guangsheng Dong, Xinglei Wang 외 arxiv

Recent geospatial foundation models (GFMs) produce spatially extensive representations of the Earth's surface that capture rich physical and environmental patterns. Among them, the AlphaEarth Foundation (AE) represents a…

Natural Language QueriesRepresentation LearningContrastive Learning

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion

2024-12-19 · Mingdeng Cao, Chong Mou, Ziyang Yuan, Xintao Wang 외

Consistent human-centric image and video synthesis aims to generate images or videos with new poses while preserving appearance consistency with a given reference image, which is crucial for low-cost visual content creat…

Computational EfficiencyDenoisingVideo Generation