paper-with-me

Papers

KSDiff: Keyframe-Augmented Speech-Aware Dual-Path Diffusion for Facial Animation

2025-09-24 · Tianle Lyu, Junchuan Zhao, Ye Wang arxiv

Audio-driven facial animation has made significant progress in multimedia applications, with diffusion models showing strong potential for talking-face synthesis. However, most existing works treat speech features as a monolithic representation and fail to capture their fine-grained roles in driving different facial motions, while also overlooking the importance of modeling keyframes with intense dynamics. To address these limitations, we propose KSDiff, a Keyframe-Augmented Speech-Aware Dual-Path Diffusion framework. Specifically, the raw audio and transcript are processed by a Dual-Path Speech Encoder (DPSE) to disentangle expression-related and head-pose-related features, while an autoregressive Keyframe Establishment Learning (KEL) module predicts the most salient motion frames. These components are integrated into a Dual-path Motion generator to synthesize coherent and realistic facial motions. Extensive experiments on HDTF and VoxCeleb demonstrate that KSDiff achieves state-of-the-art performance, with improvements in both lip synchronization accuracy and head-pose naturalness. Our results highlight the effectiveness of combining speech disentanglement with keyframe-aware diffusion for talking-head generation. The demo page is available at: https://kincin.github.io/KSDiff/.

📄 PDF Abstract BibTeX arXiv:2509.20128

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Kernel Space Diffusion Model for Efficient Remote Sensing Pansharpening

2025-05-25 · Hancong Jin, ZiHan Cao, LiangJian Deng

Pansharpening is a fundamental task in remote sensing that integrates high-resolution panchromatic imagery (PAN) with low-resolution multispectral imagery (LRMS) to produce an enhanced image with both high spatial and sp…

Pansharpening

Structure-Augmented Standard Plane Detection with Temporal Aggregation in Blind-Sweep Fetal Ultrasound

2026-04-22 · Keli Niu, He Zhao, Qianhui Men arxiv

In low-resource settings, blind-sweep ultrasound provides a practical and accessible method for identifying fetal growth restriction. However, unlike freehand ultrasound which is subjectively controlled, detection of bio…

Real-Time Synchronized Interaction Framework for Emotion-Aware Humanoid Robots

2026-01-24 · Yanrong Chen, Xihan Bian arxiv

As humanoid robots increasingly introduced into social scene, achieving emotionally synchronized multimodal interaction remains a significant challenges. To facilitate the further adoption and integration of humanoid rob…

KEMO: Event-Driven Keyframe Memory for Long-Horizon Robot Manipulation with VLA Policies

2026-06-22 · Yihan Zeng, Minghao Ye, Yiyuan Chen, Yide Shentu 외 arxiv

Long-horizon robot manipulation remains challenging because similar observations may occur at different execution stages, while the appropriate action depends on previously completed operations. Memory can address this a…

Robot Manipulation

End-to-end Speech Recognition with similar length speech and text

2025-10-12 · Peng Fan, Wenping Wang, Fei Deng arxiv

The mismatch of speech length and text length poses a challenge in automatic speech recognition (ASR). In previous research, various approaches have been employed to align text with speech, including the utilization of C…

Speech Recognition