paper-with-me

홈 › Papers

MindFlow: Harmonizing Cognitive Semantics and Acoustic Dynamics for Facial Animation Generation in Dyadic Conversations

2026-06-26 · Hejia Chen, Haoxian Zhang, Xu He, Xiaoqiang Liu, Pengfei Wan, Shoulong Zhang, Shuai Li arxiv

Generating lifelike facial animation for dyadic conversations requires reconciling high-level cognitive intent with precise low-level motor reflexes, yet existing methods fall short in the semantic understanding of dialogue context and in precise dynamic control. In this paper, we propose MindFlow, a dual-pathway generative framework inspired by the Ventral-Dorsal pathway model in neuroscience, which decouples generation into two collaborative streams, thereby harmonizing deep semantic reasoning with fine-grained control. In the Ventral module, we transform the conventional Sentence-Action approach into a novel Chunk-State approach that models raw acoustic streams as a context-aware, evolving emotional state chain, capturing subtle paralinguistic nuances and mid-utterance emotional shifts missed by sentence-level modeling. The Dorsal module features a conditional autoregressive flow matching network for high-fidelity facial motion, driven by high-frequency acoustic cues and modulated by emotion states, plus a Selective Acoustic Injector for adaptive audio gating to ensure robustness in talking-and-listening dynamics without interference. Extensive experiments demonstrate that MindFlow achieves superior semantic appropriateness and motion naturalness compared to state-of-the-art baselines.

📄 PDF Abstract BibTeX arXiv:2606.27779

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Cognitive performance in open-plan office acoustic simulations: Effects of room acoustics and semantics but not spatial separation of sound sources

2023-06-13 · Manuj Yadav, Markus Georgi, Larissa Leist, Maria Klatte 외

The irrelevant sound effect (ISE) characterizes short-term memory performance impairment during irrelevant sounds relative to quiet. Irrelevant sound presentation in most laboratory-based ISE studies has been rather limi…

Structures Meet Semantics: Multimodal Fusion via Graph Contrastive Learning

2025-08-24 · Jiangfeng Sun, Sihao He, Zhonghong Ou, Meina Song arxiv

Multimodal sentiment analysis (MSA) aims to infer emotional states by effectively integrating textual, acoustic, and visual modalities. Despite notable progress, existing multimodal fusion methods often neglect modality-…

Multimodal Sentiment AnalysisContrastive Learning

MindFlow: Revolutionizing E-commerce Customer Support with Multimodal LLM Agents

2025-07-07 · Ming Gong, Xucheng Huang, ChengHan Yang, Xianhan Peng 외

Recent advances in large language models (LLMs) have enabled new applications in e-commerce customer service. However, their capabilities remain constrained in complex, multimodal scenarios. We present MindFlow, the firs…

Decision Making

MindFlow+: A Self-Evolving Agent for E-Commerce Customer Service

2025-07-25 · Ming Gong, Xucheng Huang, Ziheng Xu, Vijayan K. Asari arxiv

High-quality dialogue is crucial for e-commerce customer service, yet traditional intent-based systems struggle with dynamic, multi-turn interactions. We present MindFlow+, a self-evolving dialogue agent that learns doma…

Reinforcement LearningResponse Generation

Predicting Cognitive Load from Speech and Interaction Dynamics in Dyadic Conversations

2026-06-11 · Tahiya Chowdhury arxiv

Estimating cognitive load from speech has largely been studied in controlled laboratory settings, with limited understanding of its reliability in natural collaborative conversations. We investigate whether speech and in…