paper-with-me

홈 › Papers

MultiModal Action Conditioned Video Generation

2025-10-02 · Yichen Li, Antonio Torralba arxiv

Current video models fail as world model as they lack fine-graiend control. General-purpose household robots require real-time fine motor control to handle delicate tasks and urgent situations. In this work, we introduce fine-grained multimodal actions to capture such precise control. We consider senses of proprioception, kinesthesia, force haptics, and muscle activation. Such multimodal senses naturally enables fine-grained interactions that are difficult to simulate with text-conditioned generative models. To effectively simulate fine-grained multisensory actions, we develop a feature learning paradigm that aligns these modalities while preserving the unique information each modality provides. We further propose a regularization scheme to enhance causality of the action trajectory features in representing intricate interaction dynamics. Experiments show that incorporating multimodal senses improves simulation accuracy and reduces temporal drift. Extensive ablation studies and downstream applications demonstrate the effectiveness and practicality of our work.

📄 PDF Abstract BibTeX arXiv:2510.02287

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

LiveTalk: Real-Time Multimodal Interactive Video Diffusion via Improved On-Policy Distillation

2025-12-29 · Ethan Chern, Zhulin Hu, Bohao Tang, Jiadi Su 외 arxiv

Real-time video generation via diffusion is essential for building general-purpose multimodal interactive AI systems. However, the simultaneous denoising of all video frames with bidirectional attention via an iterative …

Text-to-Video Generation

HunyuanVideo-HOMA: Generic Human-Object Interaction in Multimodal Driven Human Animation

2025-06-10 · Ziyao Huang, Zixiang Zhou, Juan Cao, Yifeng Ma 외

To address key limitations in human-object interaction (HOI) video generation -- specifically the reliance on curated motion data, limited generalization to novel objects/scenarios, and restricted accessibility -- we int…

Human AnimationHuman-Object Interaction DetectionVideo Generation

UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation

2020-02-15 · Huaishao Luo, Lei Ji, Botian Shi, Haoyang Huang 외

With the recent success of the pre-training technique for NLP and image-linguistic tasks, some video-linguistic pre-training works are gradually developed to improve video-text related downstream tasks. However, most of …

Action SegmentationDecoderLanguage ModelingLanguage Modelling+2

U-Mind: A Unified Framework for Real-Time Multimodal Interaction with Audiovisual Generation

2026-02-27 · Xiang Deng, Feng Gao, Yong Zhang, Youxin Pang 외 arxiv

Full-stack multimodal interaction in real-time is a central goal in building intelligent embodied agents capable of natural, dynamic communication. However, existing systems are either limited to unimodal generation or s…

Instruction FollowingQuestion Answering

DrivingGPT: Unifying Driving World Modeling and Planning with Multi-modal Autoregressive Transformers

2024-12-24 · Yuntao Chen, Yuqi Wang, Zhaoxiang Zhang

World model-based searching and planning are widely recognized as a promising path toward human-level physical intelligence. However, current driving world models primarily rely on video diffusion models, which specializ…

NavSimTrajectory PlanningVideo Generation