paper-with-me

홈 › Papers

MOSRA: Joint Mean Opinion Score and Room Acoustics Speech Quality Assessment

2022-04-04 · Karl El Hajal, Milos Cernak, Pablo Mainar

The acoustic environment can degrade speech quality during communication (e.g., video call, remote presentation, outside voice recording), and its impact is often unknown. Objective metrics for speech quality have proven challenging to develop given the multi-dimensionality of factors that affect speech quality and the difficulty of collecting labeled data. Hypothesizing the impact of acoustics on speech quality, this paper presents MOSRA: a non-intrusive multi-dimensional speech quality metric that can predict room acoustics parameters (SNR, STI, T60, DRR, and C50) alongside the overall mean opinion score (MOS) for speech quality. By explicitly optimizing the model to learn these room acoustics parameters, we can extract more informative features and improve the generalization for the MOS task when the training data is limited. Furthermore, we also show that this joint training method enhances the blind estimation of room acoustics, improving the performance of current state-of-the-art models. An additional side-effect of this joint prediction is the improvement in the explainability of the predictions, which is a valuable feature for many applications.

📄 PDF Abstract BibTeX arXiv:2204.01345

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Multi-Channel MOSRA: Mean Opinion Score and Room Acoustics Estimation Using Simulated Data and a Teacher Model

2023-09-21 · Jozef Coldenhoff, Andrew Harper, Paul Kendrick, Tijana Stojkovic 외

Previous methods for predicting room acoustic parameters and speech quality metrics have focused on the single-channel case, where room acoustics and Mean Opinion Score (MOS) are predicted for a single recording device. …

DescriptivePrediction

EchoMark: Perceptual Acoustic Environment Transfer with Watermark-Embedded Room Impulse Response

2025-11-09 · Chenpei Huang, Lingfeng Yao, Kyu In Lee, Lan Emily Zhang 외 arxiv

Acoustic Environment Matching (AEM) is the task of transferring clean audio into a target acoustic environment, enabling engaging applications such as audio dubbing and auditory immersive virtual reality (VR). Recovering…

Pilot Study on Student Public Opinion Regarding GAI

2026-01-07 · William Franz Lamberti, Sunbin Kim, Samantha Rose Lawrence arxiv

The emergence of generative AI (GAI) has sparked diverse opinions regarding its appropriate use across various domains, including education. This pilot study investigates university students' perceptions of GAI in higher…

A Multi-task Learning Framework for Opinion Triplet Extraction

2020-10-04 · Findings of the Association for Computational Linguistics 2020 · Chen Zhang, Qiuchi Li, Dawei Song, Benyou Wang

The state-of-the-art Aspect-based Sentiment Analysis (ABSA) approaches are mainly based on either detecting aspect terms and their corresponding sentiment polarities, or co-extracting aspect and opinion terms. However, t…

Aspect-Based Sentiment AnalysisAspect-Based Sentiment Analysis (ABSA)Aspect Sentiment Triplet ExtractionExtract Aspect+3

Courtroom Analogy: New Perspective on Uncertainty-Aware Classification

2026-05-25 · Taeseong Yoon, Heeyoung Kim arxiv

Single-pass uncertainty quantification (UQ) methods for classification represent uncertainty by predicting a tractable distribution over the class probability vector. While existing approaches primarily focus on enhancin…