paper-with-me

Papers

Libra-VLA: Achieving Learning Equilibrium via Asynchronous Coarse-to-Fine Dual-System

2026-04-27 · Yifei Wei, Linqing Zhong, Yi Liu, Yuxiang Lu, Xindong He, Maoqing Yao, Guanghui Ren arxiv

Vision-Language-Action (VLA) models are a promising paradigm for generalist robotic manipulation by grounding high-level semantic instructions into executable physical actions. However, prevailing approaches typically adopt a monolithic generation paradigm, directly mapping visual-linguistic features to high-frequency motor commands in a flat, non-hierarchical fashion. This strategy overlooks the inherent hierarchy of robotic manipulation, where complex actions can be naturally modeled in a Hybrid Action Space, decomposing into discrete macro-directional reaching and continuous micro-pose alignment, severely widening the semantic-actuation gap and imposing a heavy representational burden on grounding high-level semantics to continuous actions. To address this, we introduce Libra-VLA, a novel Coarse-to-Fine Dual-System VLA architecture. We explicitly decouple the learning complexity into a coarse-to-fine hierarchy to strike a training equilibrium, while simultaneously leveraging this structural modularity to implement an asynchronous execution strategy. The Semantic Planner predicts discrete action tokens capturing macro-directional intent, while the Action Refiner conditions on coarse intent to generate high-frequency continuous actions for precise alignment. Crucially, our empirical analysis reveals that performance follows an inverted-U curve relative to action decomposition granularity, peaking exactly when the learning difficulty is balanced between the two sub-systems. With the asynchronous design, our approach offers a scalable, robust, and responsive solution for open-world manipulation.

📄 PDF Abstract BibTeX arXiv:2604.24921

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

L-SNet: from Region Localization to Scale Invariant Medical Image Segmentation

2021-02-11 · Jiahao Xie, Sheng Zhang, Jianwei Lu, Ye Luo

Coarse-to-fine models and cascade segmentation architectures are widely adopted to solve the problem of large scale variations in medical image segmentation. However, those methods have two primary limitations: the first…

Image SegmentationMedical Image SegmentationSegmentationSemantic Segmentation

Semantics Lead the Way: Harmonizing Semantic and Texture Modeling with Asynchronous Latent Diffusion

2025-12-04 · Yueming Pan, Ruoyu Feng, Qi Dai, Yuqi Wang 외 arxiv

Latent Diffusion Models (LDMs) inherently follow a coarse-to-fine generation process, where high-level semantic structure is generated slightly earlier than fine-grained texture. This indicates the preceding semantics po…

FedPSA: Modeling Behavioral Staleness in Asynchronous Federated Learning

2026-02-17 · Chaoyi Lu, Yiding Sun, Zhichuan Yang, Jinqian Chen 외 arxiv

Asynchronous Federated Learning (AFL) has emerged as a significant research area in recent years. By not waiting for slower clients and executing the training process concurrently, it achieves faster training speed compa…

Federated Learning

Calibration of a neural network ocean closure for improved mean state and variability

2026-04-07 · Pavel Perezhogin, Alistair Adcroft, Laure Zanna arxiv

Global ocean models exhibit biases in the mean state and variability, particularly at coarse resolution, where mesoscale eddies are unresolved. To address these biases, parameterization coefficients are typically tuned a…

Enhanced Sampling for Efficient Learning of Coarse-Grained Machine Learning Potentials

2025-10-13 · Weilong Chen, Franz Görlich, Paul Fuchs, Julija Zavadlav arxiv

Coarse-graining (CG) enables molecular dynamics (MD) simulations of larger systems and longer timescales that are otherwise infeasible with atomistic models. Machine learning potentials (MLPs), with their capacity to cap…