paper-with-me

Papers

Before the Body Moves: Learning Anticipatory Joint Intent for Language-Conditioned Humanoid Control

2026-05-14 · Haozhe Jia, Honglei Jin, Yuan Zhang, Youcheng Fan, Shaofeng Liang, Lei Wang, Shuxu Jin, Kuimou Yu, Zinuo Zhang, Jianfei Song, Wenshuo Chen, Yutao Yue arxiv

Natural language is an intuitive interface for humanoid robots, yet streaming whole-body control requires control representations that are executable now and anticipatory of future physical transitions. Existing language-conditioned humanoid systems typically generate kinematic references that a low-level tracker must repair reactively, or use latent/action policies whose outputs do not explicitly encode upcoming contact changes, support transfers, and balance preparation. We propose \textbf{DAJI} (\emph{Dynamics-Aligned Joint Intent}), a hierarchical framework that learns an anticipatory joint-intent interface between language generation and closed-loop control. DAJI-Act distills a future-aware teacher into a deployable diffusion action policy through student-driven rollouts, while DAJI-Flow autoregressively generates future intent chunks from language and intent history. Experiments show that DAJI achieves strong results in anticipatory latent learning, single-instruction generation, and streaming instruction following, reaching 94.42\% rollout success on HumanML3D-style generation and 0.152 subsequence FID on BABEL.

📄 PDF Abstract BibTeX arXiv:2605.14417

Code (0)

등록된 구현이 없습니다.

Tasks

Instruction Following

Similar Papers 제목 키워드 기반

EgoIntent: An Egocentric Step-level Benchmark for Understanding What, Why, and Next

2026-03-12 · Ye Pan, Chi Kit Wong, Yuanhuiyi Lyu, Hanqian Li 외 arxiv

Multimodal Large Language Models (MLLMs) have demonstrated remarkable video reasoning capabilities across diverse tasks. However, their ability to understand human intent at a fine-grained level in egocentric videos rema…

AI-Instruments: Embodying Prompts as Instruments to Abstract & Reflect Graphical Interface Commands as General-Purpose Tools

2025-02-26 · Nathalie Riche, Anna Offenwanger, Frederic Gmeiner, David Brown 외

Chat-based prompts respond with verbose linear-sequential texts, making it difficult to explore and refine ambiguous intents, back up and reinterpret, or shift directions in creative AI-assisted design work. AI-Instrumen…

Image Generation

BiPreManip: Learning Affordance-Based Bimanual Preparatory Manipulation through Anticipatory Collaboration

2026-03-23 · Yan Shen, Feng Jiang, Zichen He, Xiaoqi Li 외 arxiv

Many everyday objects are difficult to directly grasp (e.g., a flat iPad) or manipulate functionally (e.g., opening the cap of a pen lying on a desk). Such tasks require sequential, asymmetric coordination between two ar…

AHEAD: Anticipatory Hand-Driven Teleoperation via Human Intent Prediction

2026-07-16 · Seok Joon Kim, Junho Lee, Federica Spinola, Taein Kwon 외 arxiv

Direct hand-driven teleoperation maps an operator's hand motion to robot end-effector commands at every frame, enabling precise control, but it requires constant monitoring and correction during approach, grasp, and plac…

TeleGate: Whole-Body Humanoid Teleoperation via Gated Expert Selection with Motion Prior

2026-02-10 · Jie Li, Bing Tang, Feng Wu arxiv

Real-time whole-body teleoperation is a critical method for humanoid robots to perform complex tasks in unstructured environments. However, developing a unified controller that robustly supports diverse human motions rem…

Knowledge Distillation