paper-with-me

홈 › Papers

ACG: Action Coherence Guidance for Flow-based Vision-Language-Action models

2025-10-25 · Minho Park, Kinam Kim, Junha Hyung, Hyojin Jang, Hoiyeong Jin, Jooyeol Yun, Hojoon Lee, Jaegul Choo arxiv

Diffusion and flow matching models have emerged as powerful robot policies, enabling Vision-Language-Action (VLA) models to generalize across diverse scenes and instructions. Yet, when trained via imitation learning, their high generative capacity makes them sensitive to noise in human demonstrations: jerks, pauses, and jitter which reduce action coherence. Reduced action coherence causes instability and trajectory drift during deployment, failures that are catastrophic in fine-grained manipulation where precision is crucial. In this paper, we present Action Coherence Guidance (ACG) for VLA models, a training-free test-time guidance algorithm that improves action coherence and thereby yields performance gains. Evaluated on RoboCasa, DexMimicGen, and real-world SO-101 tasks, ACG consistently improves action coherence and boosts success rates across diverse manipulation tasks. Code and project page are available at https://github.com/DAVIAN-Robotics/ACG and https://DAVIAN-Robotics.github.io/ACG , respectively.

📄 PDF Abstract BibTeX arXiv:2510.22201

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Guided Action Flow: Q-Guided Inference for Flow-Matching Vision-Language-Action Policies

2026-07-02 · Liuhaichen Yang, Zhuang Jiang, Chenchao Sheng, Zezhi Tang arxiv

Flow-matching vision-language-action policies generate robot action chunks through an iterative transport process, creating an opportunity for test-time guidance without retraining the base policy. We study this opportun…

Co-Annotator: Expert-Distilled ViT and VLM for Visual and Documentation Guidance in Age-Related Macular Degeneration

2026-08-31 · Ziheng "Leo" Li, Benjamin Freeman, Akshay Raman, Kavin Aravindhan Rajkumar 외 arxiv

Clinical AI often optimizes predictive performance without engaging how clinicians decide where to look and what to write. We present Co-Annotator, which distills expert gaze and dictation into two guidance components: a…

Routing Matters in MoE: Scaling Diffusion Transformers with Explicit Routing Guidance

2025-10-28 · Yujie Wei, Shiwei Zhang, Hangjie Yuan, Yujin Han 외 arxiv

Mixture-of-Experts (MoE) has emerged as a powerful paradigm for scaling model capacity while preserving computational efficiency. Despite its notable success in large language models (LLMs), existing attempts to apply Mo…

Computational Efficiency

Neuro-Symbolic Safety Guidance for Vision-Language-Action Models via Constrained Flow Matching

2026-07-01 · William English, Hao Zheng, Rickard Ewetz arxiv

Vision-Language-Action (VLA) models have demonstrated promising generalization capabilities across robotic manipulation tasks, yet their real-world deployment remains limited by the lack of effective safety measures. Spe…

Collision Avoidance

Eulerian Motion Guidance: Robust Image Animation via Bidirectional Geometric Consistency

2026-05-07 · Thong Nguyen, Khoi M. Le, Cong-Duy Nguyen, Luu Anh Tuan 외 arxiv

Recent advancements in image animation have utilized diffusion models to breathe life into static images. However, existing controllable frameworks typically rely on Lagrangian motion guidance, where optical flow is esti…