paper-with-me

홈 › Papers

GeoAlign: Beyond Semantics with State-Guided Spatial Alignment in VLA Models

2026-06-02 · Yizhi Chen, Zhanxiang Cao, Xinyi Peng, Yixiao Zheng, Xiaxi Si, Yiheng Li, Liyun Yan, Keqi Zhu, Xueyun Chen, Shengcheng Fu, Tianyue Zhan, Yufei Jia, Jinming Yao, Yan Xie, Kun Wang, Cewu Lu, Yue Gao arxiv

Current Vision--Language--Action (VLA) models often optimize for semantic grounding, whereas executable manipulation requires geometry-aware spatial alignment and dynamic affordance selection. We introduce GeoAlign, a state-guided spatial alignment architecture for VLA policy learning. GeoAlign post-trains an RGB geometry branch with robot-domain RGB-D supervision, yielding RGB-derived Geometry-Enhanced Post-Trained (GEP) features for policy rollout. The robot's proprioceptive state queries the GEP feature grid, producing compact, phase-dependent geometry tokens for action prediction. GeoAlign achieves 99.0% on LIBERO, 85.3% across three SimplerEnv-Fractal tasks, and 78.8% on eight geometry-critical real-world ALOHA tasks, with ablations confirming the value of geometry post-training and proprioceptive-state-guided querying.

📄 PDF Abstract BibTeX arXiv:2606.03240

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

GeoAlign: Geometric Feature Realignment for MLLM Spatial Reasoning

2026-04-14 · Zhaochen Liu, Limeng Qiao, Guanglu Wan, Tingting Jiang arxiv

Multimodal large language models (MLLMs) have exhibited remarkable performance in various visual tasks, yet still struggle with spatial reasoning. Recent efforts mitigate this by injecting geometric features from 3D foun…

Spatial Reasoning

GEOALIGN: Geometric Rollout Curation for Robust LLM Reinforcement Learning

2026-06-25 · Ting Zhou, Zhenqing Ling, Yiyang Zhao, Ying Shen 외 arxiv

Online reinforcement learning is widely used to align large language models (LLMs) with reward signals, yet training can be unstable under noisy or misspecified rewards. We identify a failure mode we call directional inc…

Reinforcement LearningMathematical Reasoning

All You Need is a Second Look: Towards Arbitrary-Shaped Text Detection

2021-06-24 · Meng Cao, Can Zhang, Dongming Yang, Yuexian Zou

Arbitrary-shaped text detection is a challenging task since curved texts in the wild are of the complex geometric layouts. Existing mainstream methods follow the instance segmentation pipeline to obtain the text regions.…

AllInstance SegmentationSegmentationSemantic Segmentation+1

GeoAlignCLIP: Enhancing Fine-Grained Vision-Language Alignment in Remote Sensing via Multi-Granular Consistency Learning

2026-03-10 · Xiao Yang, Ronghao Fu, Zhuoran Duan, Zhiwen Lin 외 arxiv

Vision-language pretraining models have made significant progress in bridging remote sensing imagery with natural language. However, existing approaches often fail to effectively integrate multi-granular visual and textu…

Beyond AlphaEarth: Toward Human-Centered Geospatial Foundation Models via POI-Guided Contrastive Learning

2025-10-10 · Junyuan Liu, Quan Qin, Guangsheng Dong, Xinglei Wang 외 arxiv

Recent geospatial foundation models (GFMs) produce spatially extensive representations of the Earth's surface that capture rich physical and environmental patterns. Among them, the AlphaEarth Foundation (AE) represents a…

Natural Language QueriesRepresentation LearningContrastive Learning