paper-with-me

홈 › Papers

FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation

2025-08-15 · MengChao Wang, Qiang Wang, Fan Jiang, Mu Xu arxiv

Recent advances in audio-driven portrait animation have demonstrated impressive capabilities. However, existing methods struggle to align with fine-grained human preferences across multiple dimensions, such as motion naturalness, lip-sync accuracy, and visual quality. This is due to the difficulty of optimizing among competing preference objectives, which often conflict with one another, and the scarcity of large-scale, high-quality datasets with multidimensional preference annotations. To address these, we first introduce Talking-Critic, a multimodal reward model that learns human-aligned reward functions to quantify how well generated videos satisfy multidimensional expectations. Leveraging this model, we curate Talking-NSQ, a large-scale multidimensional human preference dataset containing 410K preference pairs. Finally, we propose Timestep-Layer adaptive multi-expert Preference Optimization (TLPO), a novel framework for aligning diffusion-based portrait animation models with fine-grained, multidimensional preferences. TLPO decouples preferences into specialized expert modules, which are then fused across timesteps and network layers, enabling comprehensive, fine-grained enhancement across all dimensions without mutual interference. Experiments demonstrate that Talking-Critic significantly outperforms existing methods in aligning with human preference ratings. Meanwhile, TLPO achieves substantial improvements over baseline models in lip-sync accuracy, motion naturalness, and visual quality, exhibiting superior performance in both qualitative and quantitative evaluations. Ours project page: https://fantasy-amap.github.io/fantasy-talking2/

📄 PDF Abstract BibTeX arXiv:2508.11255

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Spiking Layer-Adaptive Magnitude-based Pruning

2026-03-16 · Junqiao Wang, Zhehang Ye, Yuqi Ouyang arxiv

Spiking Neural Networks (SNNs) provide energy-efficient computation but their deployment is constrained by dense connectivity and high spiking operation costs. Existing magnitude-based pruning strategies, when naively ap…

Rethinking Direct Preference Optimization in Diffusion Models

2025-05-24 · Junyong Kang, Seohyun Lim, Kyungjune Baek, Hyunjung Shim

Aligning text-to-image (T2I) diffusion models with human preferences has emerged as a critical research challenge. While recent advances in this area have extended preference optimization techniques from large language m…

Diffusion Model as a Noise-Aware Latent Reward Model for Step-Level Preference Optimization

2025-02-03 · Tao Zhang, Cheng Da, Kun Ding, Huan Yang 외

Preference optimization for diffusion models aims to align them with human preferences for images. Previous methods typically use Vision-Language Models (VLMs) as pixel-level reward models to approximate human preference…

model

Nucleus-Image: Sparse MoE for Image Generation

2026-04-14 · Chandan Akiti, Ajay Modukuri, Murali Nandan Nagarapu, Gunavardhan Akiti 외 arxiv

We present Nucleus-Image, a text-to-image generation model that establishes a new Pareto frontier in quality-versus-efficiency by matching or exceeding leading models on GenEval, DPG-Bench, and OneIG-Bench while activati…

Text-to-Image GenerationReinforcement Learning

DS-ATGO: Dual-Stage Synergistic Learning via Forward Adaptive Threshold and Backward Gradient Optimization for Spiking Neural Networks

2025-11-17 · Jiaqiang Jiang, Wenfeng Xu, Jing Fan, Rui Yan arxiv

Brain-inspired spiking neural networks (SNNs) are recognized as a promising avenue for achieving efficient, low-energy neuromorphic computing. Direct training of SNNs typically relies on surrogate gradient (SG) learning …