paper-with-me

홈 › Papers

EvalTalker: Learning to Evaluate Real-Portrait-Driven Multi-Subject Talking Humans

2025-12-01 · Yingjie Zhou, Xilei Zhu, Siyu Ren, Ziyi Zhao, Ziwen Wang, Farong Wen, Yu Zhou, Jiezhang Cao, Xiongkuo Min, Fengjiao Chen, Xiaoyu Li, Xuezhi Cao, Guangtao Zhai, Xiaohong Liu arxiv

Speech-driven Talking Human (TH) generation, commonly known as "Talker," currently faces limitations in multi-subject driving capabilities. Extending this paradigm to "Multi-Talker," capable of animating multiple subjects simultaneously, introduces richer interactivity and stronger immersion in audiovisual communication. However, current Multi-Talkers still exhibit noticeable quality degradation caused by technical limitations, resulting in suboptimal user experiences. To address this challenge, we construct THQA-MT, the first large-scale Multi-Talker-generated Talking Human Quality Assessment dataset, consisting of 5,492 Multi-Talker-generated THs (MTHs) from 15 representative Multi-Talkers using 400 real portraits collected online. Through subjective experiments, we analyze perceptual discrepancies among different Multi-Talkers and identify 12 common types of distortion. Furthermore, we introduce EvalTalker, a novel TH quality assessment framework. This framework possesses the ability to perceive global quality, human characteristics, and identity consistency, while integrating Qwen-Sync to perceive multimodal synchrony. Experimental results demonstrate that EvalTalker achieves superior correlation with subjective scores, providing a robust foundation for future research on high-quality Multi-Talker generation and evaluation.

📄 PDF Abstract BibTeX arXiv:2512.01340

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FlowPortrait: Reinforcement Learning for Audio-Driven Portrait Video Generation

2026-02-25 · Weiting Tan, Andy T. Liu, Ming Tu, Xinghua Qu 외 arxiv

Generating realistic talking-head videos remains challenging due to persistent issues such as imperfect lip synchronization, unnatural motion, and evaluation metrics that correlate poorly with human perception. We propos…

Reinforcement LearningVideo Generation

MODA: Mapping-Once Audio-driven Portrait Animation with Dual Attentions

2023-07-19 · ICCV 2023 1 · Yunfei Liu, Lijian Lin, Fei Yu, Changyin Zhou 외

Audio-driven portrait animation aims to synthesize portrait videos that are conditioned by given audio. Animating high-fidelity and multimodal video portraits has a variety of applications. Previous methods have attempte…

Portrait Animation

AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

2024-03-26 · Huawei Wei, Zejun Yang, Zhisheng Wang

In this study, we propose AniPortrait, a novel framework for generating high-quality animation driven by audio and a reference portrait image. Our methodology is divided into two stages. Initially, we extract 3D intermed…

DiversityFace ReenactmentPortrait Animation

FREAK: Frequency-modulated High-fidelity and Real-time Audio-driven Talking Portrait Synthesis

2025-03-06 · Ziqi Ni, Ao Fu, Yi Zhou

Achieving high-fidelity lip-speech synchronization in audio-driven talking portrait synthesis remains challenging. While multi-stage pipelines or diffusion models yield high-quality results, they suffer from high computa…

Audio-Visual Synchronization

SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation

2026-04-09 · Wenli Zhang, Xianglong Shi, Sirui Zhao, Xinqi Chen 외 arxiv

Diffusion-based audio-driven talking-head generation enables realistic portrait animation, but also introduces risks of misuse, such as fraud and misinformation. Existing protection methods are largely limited to a singl…

Talking Head Generation