paper-with-me

Papers

Towards Reliable Objective Evaluation Metrics for Generative Singing Voice Separation Models

2025-07-15 · Paul A. Bereuter, Benjamin Stahl, Mark D. Plumbley, Alois Sontacchi

Traditional Blind Source Separation Evaluation (BSS-Eval) metrics were originally designed to evaluate linear audio source separation models based on methods such as time-frequency masking. However, recent generative models may introduce nonlinear relationships between the separated and reference signals, limiting the reliability of these metrics for objective evaluation. To address this issue, we conduct a Degradation Category Rating listening test and analyze correlations between the obtained degradation mean opinion scores (DMOS) and a set of objective audio quality metrics for the task of singing voice separation. We evaluate three state-of-the-art discriminative models and two new competitive generative models. For both discriminative and generative models, intrusive embedding-based metrics show higher correlations with DMOS than conventional intrusive metrics such as BSS-Eval. For discriminative models, the highest correlation is achieved by the MSE computed on Music2Latent embeddings. When it comes to the evaluation of generative models, the strongest correlations are evident for the multi-resolution STFT loss and the MSE calculated on MERT-L12 embeddings, with the latter also providing the most balanced correlation across both model types. Our results highlight the limitations of BSS-Eval metrics for evaluating generative singing voice separation models and emphasize the need for careful selection and validation of alternative evaluation metrics for the task of singing voice separation.

📄 PDF Abstract BibTeX arXiv:2507.11427

Code (1)

pablebe/gensvs_eval 공식 구현 pytorch

Tasks

Audio Source Separationblind source separation

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Score and Lyrics-Free Singing Voice Generation

2019-12-26 · Jen-Yu Liu, Yu-Hua Chen, Yin-Cheng Yeh, Yi-Hsuan Yang

Generative models for singing voice have been mostly concerned with the task of ``singing voice synthesis,'' i.e., to produce singing voice waveforms given musical scores and text lyrics. In this work, we explore a novel…

Audio GenerationSinging Voice Synthesis

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model

2024-12-04 · Yan Li, Ziya Zhou, Zhiqiang Wang, Wei Xue 외

Recent advancements in generative models have significantly enhanced talking face video generation, yet singing video generation remains underexplored. The differences between human talking and singing limit the performa…

Video Generation

Learning the Beauty in Songs: Neural Singing Voice Beautifier

2022-02-27 · ACL 2022 5 · Jinglin Liu, Chengxi Li, Yi Ren, Zhiying Zhu 외

We are interested in a novel task, singing voice beautifying (SVB). Given the singing voice of an amateur singer, SVB aims to improve the intonation and vocal tone of the voice, while keeping the content and vocal timbre…

Dynamic Time Warping

AnyEnhance: A Unified Generative Model with Prompt-Guidance and Self-Critic for Voice Enhancement

2025-01-26 · Junan Zhang, Jing Yang, Zihao Fang, Yuancheng Wang 외

We introduce AnyEnhance, a unified generative model for voice enhancement that processes both speech and singing voices. Based on a masked generative model, AnyEnhance is capable of handling both speech and singing voice…

DenoisingIn-Context LearningSuper-ResolutionTarget Speaker Extraction

SingMOS-Pro: An Comprehensive Benchmark for Singing Quality Assessment

2025-10-02 · Yuxun Tang, Lan Liu, Wenhao Feng, Yiwen Zhao 외 arxiv

Singing voice generation progresses rapidly, yet evaluating singing quality remains a critical challenge. Human subjective assessment, typically in the form of listening tests, is costly and time consuming, while existin…