paper-with-me

홈 › Papers

Adapting Frechet Audio Distance for Generative Music Evaluation

2023-11-02 · Azalea Gui, Hannes Gamper, Sebastian Braun, Dimitra Emmanouilidou

The growing popularity of generative music models underlines the need for perceptually relevant, objective music quality metrics. The Frechet Audio Distance (FAD) is commonly used for this purpose even though its correlation with perceptual quality is understudied. We show that FAD performance may be hampered by sample size bias, poor choice of audio embeddings, or the use of biased or low-quality reference sets. We propose reducing sample size bias by extrapolating scores towards an infinite sample size. Through comparisons with MusicCaps labels and a listening test we identify audio embeddings and music reference sets that yield FAD scores well-correlated with acoustic and musical quality. Our results suggest that per-song FAD can be useful to identify outlier samples and predict perceptual quality for a range of music sets and generative models. Finally, we release a toolkit that allows adapting FAD for generative music evaluation.

📄 PDF Abstract BibTeX arXiv:2311.01616

Code (5)

microsoft/fadtk 공식 구현 pytorch
YoonjinXD/kadtk pytorch
dcase2024-task7-sound-scene-synthesis/fadtk pytorch
pablebe/gensvs_eval pytorch
soham97/pam pytorch

Tasks

FAD

Similar Papers 제목 키워드 기반

Frechet Music Distance: A Metric For Generative Symbolic Music Evaluation

2024-12-10 · Jan Retkowski, Jakub Stępniak, Mateusz Modrzejewski

In this paper we introduce the Frechet Music Distance (FMD), a novel evaluation metric for generative symbolic music models, inspired by the Frechet Inception Distance (FID) in computer vision and Frechet Audio Distance …

FADMusic GenerationMusic Modeling

Addressing Emotion Bias in Music Emotion Recognition and Generation with Frechet Audio Distance

2024-09-23 · Yuanchao Li, Azalea Gui, Dimitra Emmanouilidou, Hannes Gamper

The complex nature of musical emotion introduces inherent bias in both recognition and generation, particularly when relying on a single audio encoder, emotion classifier, or evaluation metric. In this work, we conduct a…

Emotion RecognitionFADMusic Emotion RecognitionMusic Generation

Can Synthetic Audio From Generative Foundation Models Assist Audio Recognition and Speech Modeling?

2024-06-13 · Tiantian Feng, Dimitrios Dimitriadis, Shrikanth Narayanan

Recent advances in foundation models have enabled audio-generative models that produce high-fidelity sounds associated with music, events, and human actions. Despite the success achieved in modern audio-generative models…

Audio GenerationData Augmentation

DOSE : Drum One-Shot Extraction from Music Mixture

2025-04-25 · Suntae Hwang, SeongHyeon Kang, KyungSu Kim, Semin Ahn 외

Drum one-shot samples are crucial for music production, particularly in sound design and electronic music. This paper introduces Drum One-Shot Extraction, a task in which the goal is to extract drum one-shots that are pr…

FAD

Art2Music: Generating Music for Art Images with Multi-modal Feeling Alignment

2025-11-27 · Jiaying Hong, Ting Zhu, Thanet Markchom, Huizhi Liang arxiv

With the rise of AI-generated content (AIGC), generating perceptually natural and feeling-aligned music from multimodal inputs has become a central challenge. Existing approaches often rely on explicit emotion labels tha…

Audio GenerationMusic Generation