paper-with-me

Papers

RF-GML: Reference-Free Generative Machine Listener

2024-09-16 · Arijit Biswas, Guanxin Jiang

This paper introduces a novel reference-free (RF) audio quality metric called the RF-Generative Machine Listener (RF-GML), designed to evaluate coded mono, stereo, and binaural audio at a 48 kHz sample rate. RF-GML leverages transfer learning from a state-of-the-art full-reference (FR) Generative Machine Listener (GML) with minimal architectural modifications. The term "generative" refers to the model's ability to generate an arbitrary number of simulated listening scores. Unlike existing RF models, RF-GML accurately predicts subjective quality scores across diverse content types and codecs. Extensive evaluations demonstrate its superiority in rating unencoded audio and distinguishing different levels of coding artifacts. RF-GML's performance and versatility make it a valuable tool for coded audio quality assessment and monitoring in various applications, all without the need for a reference signal.

📄 PDF Abstract BibTeX arXiv:2409.10210

Code (0)

등록된 구현이 없습니다.

Tasks

Audio Quality AssessmentTransfer Learning

Similar Papers 제목 키워드 기반

Generative Machine Listener

2023-08-18 · Guanxin Jiang, Lars Villemoes, Arijit Biswas

We show how a neural network can be trained on individual intrusive listening test scores to predict a distribution of scores for each pair of reference and coded input stereo or binaural signals. We nickname this method…

Data Augmentation

Listener Model for the PhotoBook Referential Game with CLIPScores as Implicit Reference Chain

2023-06-16 · Shih-Lun Wu, Yi-Hui Chou, Liangze Li

PhotoBook is a collaborative dialogue game where two players receive private, partially-overlapping sets of images and resolve which images they have in common. It presents machines with a great challenge to learn how pe…

Listener-Rewarded Thinking in VLMs for Image Preferences

2025-06-28 · Alexander Gambashidze, Li Pengyi, Matvey Skripkin, Andrey Galichin 외

Training robust and generalizable reward models for human visual preferences is essential for aligning text-to-image and text-to-video generative models with human intent. However, current reward models often fail to gen…

MemorizationReinforcement Learning (RL)

ReactMotion: Generating Reactive Listener Motions from Speaker Utterance

2026-03-16 · Cheng Luo, Bizhu Wu, Bing Li, Jianfeng Ren 외 arxiv

In this paper, we introduce a new task, Reactive Listener Motion Generation from Speaker Utterance, which aims to generate naturalistic listener body motions that appropriately respond to a speaker's utterance. However, …

Enhanced Generative Machine Listener

2025-09-25 · Vishnu Raj, Gouthaman KV, Shiv Gehlot, Lars Villemoes 외 arxiv

We present GMLv2, a reference-based model designed for the prediction of subjective audio quality as measured by MUSHRA scores. GMLv2 introduces a Beta distribution-based loss to model the listener ratings and incorporat…