paper-with-me

Papers

A Comparative Study of Perceptual Quality Metrics for Audio-driven Talking Head Videos

2024-03-11 · Weixia Zhang, Chengguang Zhu, Jingnan Gao, Yichao Yan, Guangtao Zhai, Xiaokang Yang

The rapid advancement of Artificial Intelligence Generated Content (AIGC) technology has propelled audio-driven talking head generation, gaining considerable research attention for practical applications. However, performance evaluation research lags behind the development of talking head generation techniques. Existing literature relies on heuristic quantitative metrics without human validation, hindering accurate progress assessment. To address this gap, we collect talking head videos generated from four generative methods and conduct controlled psychophysical experiments on visual quality, lip-audio synchronization, and head movement naturalness. Our experiments validate consistency between model predictions and human annotations, identifying metrics that align better with human opinions than widely-used measures. We believe our work will facilitate performance evaluation and model development, providing insights into AIGC in a broader context. Code and data will be made available at https://github.com/zwx8981/ADTH-QA.

📄 PDF Abstract BibTeX arXiv:2403.06421

Code (1)

zwx8981/adth-qa 공식 구현 pytorch

Tasks

Talking Head Generation

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems

2025-04-30 · huan zhang, Jinhua Liang, Huy Phan, Wenwu Wang 외

Evaluating generative models remains a fundamental challenge, particularly when the goal is to reflect human preferences. In this paper, we use music generation as a case study to investigate the gap between automatic ev…

Music Generation

Evaluating generative audio systems and their metrics

2022-08-31 · Ashvala Vinay, Alexander Lerch

Recent years have seen considerable advances in audio synthesis with deep generative models. However, the state-of-the-art is very difficult to quantify; different studies often use different evaluation methodologies and…

Audio Synthesis

What You Hear Is What You See: Audio Quality Metrics From Image Quality Metrics

2023-05-19 · Tashi Namgyal, Alexander Hepburn, Raul Santos-Rodriguez, Valero Laparra 외

In this study, we investigate the feasibility of utilizing state-of-the-art image perceptual metrics for evaluating audio signals by representing them as spectrograms. The encouraging outcome of the proposed approach is …

A Comparative Evaluation of Deep Learning Models for Speech Enhancement in Real-World Noisy Environments

2025-06-17 · Md Jahangir Alam Khondkar, Ajan Ahmed, Masudul Haider Imtiaz, Stephanie Schuckers

Speech enhancement, particularly denoising, is vital in improving the intelligibility and quality of speech signals for real-world applications, especially in noisy environments. While prior research has introduced vario…

DenoisingSpeaker RecognitionSpeaker VerificationSpeech Enhancement

Data is Overrated: Perceptual Metrics Can Lead Learning in the Absence of Training Data

2023-12-06 · Tashi Namgyal, Alexander Hepburn, Raul Santos-Rodriguez, Valero Laparra 외

Perceptual metrics are traditionally used to evaluate the quality of natural signals, such as images and audio. They are designed to mimic the perceptual behaviour of human observers and usually reflect structures found …