paper-with-me

Papers

Enhanced Generative Machine Listener

2025-09-25 · Vishnu Raj, Gouthaman KV, Shiv Gehlot, Lars Villemoes, Arijit Biswas arxiv

We present GMLv2, a reference-based model designed for the prediction of subjective audio quality as measured by MUSHRA scores. GMLv2 introduces a Beta distribution-based loss to model the listener ratings and incorporates additional neural audio coding (NAC) subjective datasets to extend its generalization and applicability. Extensive evaluations on diverse testset demonstrate that proposed GMLv2 consistently outperforms widely used metrics, such as PEAQ and ViSQOL, both in terms of correlation with subjective scores and in reliably predicting these scores across diverse content types and codec configurations. Consequently, GMLv2 offers a scalable and automated framework for perceptual audio quality evaluation, poised to accelerate research and development in modern audio coding technologies.

📄 PDF Abstract BibTeX arXiv:2509.21463

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RF-GML: Reference-Free Generative Machine Listener

2024-09-16 · Arijit Biswas, Guanxin Jiang

This paper introduces a novel reference-free (RF) audio quality metric called the RF-Generative Machine Listener (RF-GML), designed to evaluate coded mono, stereo, and binaural audio at a 48 kHz sample rate. RF-GML lever…

Audio Quality AssessmentTransfer Learning

Generative Machine Listener

2023-08-18 · Guanxin Jiang, Lars Villemoes, Arijit Biswas

We show how a neural network can be trained on individual intrusive listening test scores to predict a distribution of scores for each pair of reference and coded input stereo or binaural signals. We nickname this method…

Data Augmentation

Crowdsourcing MUSHRA Tests in the Age of Generative Speech Technologies: A Comparative Analysis of Subjective and Objective Testing Methods

2025-06-01 · Laura Lechler, Chamran Moradi, Ivana Balic

The MUSHRA framework is widely used for detecting subtle audio quality differences but traditionally relies on expert listeners in controlled environments, making it costly and impractical for model development. As a res…

ReactMotion: Generating Reactive Listener Motions from Speaker Utterance

2026-03-16 · Cheng Luo, Bizhu Wu, Bing Li, Jianfeng Ren 외 arxiv

In this paper, we introduce a new task, Reactive Listener Motion Generation from Speaker Utterance, which aims to generate naturalistic listener body motions that appropriately respond to a speaker's utterance. However, …

Speaker discrimination in humans and machines: Effects of speaking style variability

2020-08-08 · Amber Afshan, Jody Kreiman, Abeer Alwan

Does speaking style variation affect humans' ability to distinguish individuals from their voices? How do humans compare with automatic systems designed to discriminate between voices? In this paper, we attempt to answer…

Speaker Verification