paper-with-me

홈 › Papers

Unified Multi-Modal Interactive & Reactive 3D Motion Generation via Rectified Flow

2025-09-28 · Prerit Gupta, Shourya Verma, Ananth Grama, Aniket Bera arxiv

Generating realistic, context-aware two-person motion conditioned on diverse modalities remains a fundamental challenge for graphics, animation and embodied AI systems. Real-world applications such as VR/AR companions, social robotics and game agents require models capable of producing coordinated interpersonal behaviour while flexibly switching between interactive and reactive generation. We introduce DualFlow, the first unified and efficient framework for multi-modal two-person motion generation. DualFlow conditions 3D motion generation on diverse inputs, including text, music, and prior motion sequences. Leveraging rectified flow, it achieves deterministic straight-line sampling paths between noise and data, reducing inference time and mitigating error accumulation common in diffusion-based models. To enhance semantic grounding, DualFlow employs a novel Retrieval-Augmented Generation (RAG) module for two-person motion that retrieves motion exemplars using music features and LLM-based text decompositions of spatial relations, body movements, and rhythmic patterns. We use a contrastive rectified flow objective to further sharpen alignment with conditioning signals and add synchronisation loss to improve inter-person temporal coordination. Extensive evaluations across interactive, reactive, and multi-modal benchmarks demonstrate that DualFlow consistently improves motion quality, responsiveness, and semantic fidelity. DualFlow achieves state-of-the-art performance in two-person multi-modal motion generation, producing coherent, expressive, and rhythmically synchronized motion.

📄 PDF Abstract BibTeX arXiv:2509.24099

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CoPlanner: An Interactive Motion Planner with Contingency-Aware Diffusion for Autonomous Driving

2025-09-21 · Ruiguo Zhong, Ruoyu Yao, Pei Liu, Xiaolong Chen 외 arxiv

Accurate trajectory prediction and motion planning are crucial for autonomous driving systems to navigate safely in complex, interactive environments characterized by multimodal uncertainties. However, current generation…

Trajectory PredictionAutonomous DrivingMotion Planning

ROMA: Real-time Omni-Multimodal Assistant with Interactive Streaming Understanding

2026-01-15 · Xueyun Tian, Wei Li, Bingbing Xu, Heng Dong 외 arxiv

Recent Omni-multimodal Large Language Models show promise in unified audio, vision, and text modeling. However, streaming audio-video understanding remains challenging, as existing approaches suffer from disjointed capab…

Interactive multi-modal motion planning with Branch Model Predictive Control

2021-09-10 · Yuxiao Chen, Ugo Rosolia, Wyatt Ubellacker, Noel Csomay-Shanklin 외

Motion planning for autonomous robots and vehicles in presence of uncontrolled agents remains a challenging problem as the reactive behaviors of the uncontrolled agents must be considered. Since the uncontrolled agents u…

Autonomous VehiclesModel Predictive ControlMotion Planning

ProAct: A Dual-System Framework for Proactive Embodied Social Agents

2026-02-15 · Zeyi Zhang, Zixi Kang, Ruijie Zhao, Yusen Feng 외 arxiv

Embodied social agents have recently advanced in generating synchronized speech and gestures. However, most interactive systems remain fundamentally reactive, responding only to current sensory inputs within a short temp…

Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural Conversation

2026-01-02 · Taekyung Ki, Sangwon Jang, Jaehyeong Jo, Jaehong Yoon 외 arxiv

Talking head generation creates lifelike avatars from static portraits for virtual communication and content creation. However, current models do not yet convey the feeling of truly interactive communication, often gener…

Talking Head Generation