paper-with-me

Papers

Evaluating generative audio systems and their metrics

2022-08-31 · Ashvala Vinay, Alexander Lerch

Recent years have seen considerable advances in audio synthesis with deep generative models. However, the state-of-the-art is very difficult to quantify; different studies often use different evaluation methodologies and different metrics when reporting results, making a direct comparison to other systems difficult if not impossible. Furthermore, the perceptual relevance and meaning of the reported metrics in most cases unknown, prohibiting any conclusive insights with respect to practical usability and audio quality. This paper presents a study that investigates state-of-the-art approaches side-by-side with (i) a set of previously proposed objective metrics for audio reconstruction, and with (ii) a listening study. The results indicate that currently used objective metrics are insufficient to describe the perceptual quality of current systems.

📄 PDF Abstract BibTeX arXiv:2209.00130

Code (0)

등록된 구현이 없습니다.

Tasks

Audio Synthesis

Similar Papers 제목 키워드 기반

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems

2025-04-30 · huan zhang, Jinhua Liang, Huy Phan, Wenwu Wang 외

Evaluating generative models remains a fundamental challenge, particularly when the goal is to reflect human preferences. In this paper, we use music generation as a case study to investigate the gap between automatic ev…

Music Generation

Isolation performance metrics for personal sound zone reproduction systems

2022-09-22 · Yue Qiao, Léo Guadagnin, Edgar Choueiri

Two isolation performance metrics, Inter-Zone Isolation (IZI) and Inter-Program Isolation (IPI), are introduced for evaluating Personal Sound Zone (PSZ) systems. Compared to the commonly-used Acoustic Contrast metric, IZ…

Aligning Text-to-Music Evaluation with Human Preferences

2025-03-20 · Yichen Huang, Zachary Novack, Koichi Saito, Jiatong Shi 외

Despite significant recent advances in generative acoustic text-to-music (TTM) modeling, robust evaluation of these models lags behind, relying in particular on the popular Fr\'echet Audio Distance (FAD). In this work, w…

FAD

PodEval: A Multimodal Evaluation Framework for Podcast Audio Generation

2025-10-01 · Yujia Xiao, Liumeng Xue, Lei He, Xinyi Chen 외 arxiv

Recently, an increasing number of multimodal (text and audio) benchmarks have emerged, primarily focusing on evaluating models' understanding capability. However, exploration into assessing generative capabilities remain…

Audio Generation

Text-to-Audio Grounding Based Novel Metric for Evaluating Audio Caption Similarity

2022-10-03 · Swapnil Bhosale, Rupayan Chakraborty, Sunil Kumar Kopparapu

Automatic Audio Captioning (AAC) refers to the task of translating an audio sample into a natural language (NL) text that describes the audio events, source of the events and their relationships. Unlike NL text generatio…

Audio captioningImage CaptioningTAGText Generation