paper-with-me

Papers

Primary-Fine Decoupling for Action Generation in Robotic Imitation

2026-02-25 · Xiaohan Lei, Min Wang, Wengang Zhou, Xingyu Lu, Houqiang Li arxiv

Multi-modal distribution in robotic manipulation action sequences poses critical challenges for imitation learning. To this end, existing approaches often model the action space as either a discrete set of tokens or a continuous, latent-variable distribution. However, both approaches present trade-offs: some methods discretize actions into tokens and therefore lose fine-grained action variations, while others generate continuous actions in a single stage tend to produce unstable mode transitions. To address these limitations, we propose Primary-Fine Decoupling for Action Generation (PF-DAG), a two-stage framework that decouples coarse action consistency from fine-grained variations. First, we compress action chunks into a small set of discrete modes, enabling a lightweight policy to select consistent coarse modes and avoid mode bouncing. Second, a mode conditioned MeanFlow policy is learned to generate high-fidelity continuous actions. Theoretically, we prove PF-DAG's two-stage design achieves a strictly lower MSE bound than single-stage generative policies. Empirically, PF-DAG outperforms state-of-the-art baselines across 56 tasks from Adroit, DexArt, and MetaWorld benchmarks. It further generalizes to real-world tactile dexterous manipulation tasks. Our work demonstrates that explicit mode-level decoupling enables both robust multi-modal modeling and reactive closed-loop control for robotic manipulation.

📄 PDF Abstract BibTeX arXiv:2602.21684

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RoboAct-CLIP: Video-Driven Pre-training of Atomic Action Understanding for Robotics

2025-04-02 · Zhiyuan Zhang, Yuxin He, Yong Sun, Junyu Shi 외

Visual Language Models (VLMs) have emerged as pivotal tools for robotic systems, enabling cross-task generalization, dynamic environmental interaction, and long-horizon planning through multimodal perception and semantic…

Action UnderstandingRepresentation Learning

Language-free Compositional Action Generation via Decoupling Refinement

2023-07-07 · Xiao Liu, Guangyi Chen, Yansong Tang, Guangrun Wang 외

Composing simple elements into complex concepts is crucial yet challenging, especially for 3D action generation. Existing methods largely rely on extensive neural language annotations to discern composable latent semanti…

Action Generation

AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model

2025-11-27 · Dian Zheng, Manyuan Zhang, Hongyu Li, Kai Zou 외 arxiv

Unified multimodal models for image generation and understanding represent a significant step toward AGI and have attracted widespread attention from researchers. The main challenge of this task lies in the difficulty in…

Image Generation

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics

2025-06-02 · Mustafa Shukor, Dana Aubakirova, Francesco Capuano, Pepijn Kooijmans 외

Vision-language models (VLMs) pretrained on large-scale multimodal datasets encode rich visual and linguistic knowledge, making them a strong foundation for robotics. Rather than training robotic policies from scratch, r…

Action GenerationGPUVision-Language-Action

Unified Video Action Model

2025-02-28 · Shuang Li, Yihuai Gao, Dorsa Sadigh, Shuran Song

A unified video and action model holds significant promise for robotics, where videos provide rich scene information for action prediction, and actions provide dynamics information for video prediction. However, effectiv…

modelPredictionVideo GenerationVideo Prediction