paper-with-me

홈 › Papers

Preference-ASR: A Preference-Aware Test Set for Benchmarking ASR in the Era of Speech LLMs

2026-06-28 · Nithin Rao Koluguri, Sasha Meister, Nikolay Karpov, Piotr Zelasko, Desh Raj, Jagadeesh Balam, Boris Ginsburg arxiv

Popular ASR test sets adopt inconsistent conventions for numbers, disfluencies, entities, and casing, while standard normalizers erase the format distinctions users care about. Current benchmarks therefore cannot measure whether a model follows user preferences for output style. We introduce PreferenceASR, a test set evaluating ASR systems on their ability to follow natural-language preference instructions across four categories: normalization, entities, disfluencies, and case. Built from seven open-source corpora via a two-stage LLM-assisted pipeline with human verification, it is evaluated with a preference-aware normalizer that selectively skips steps matching the active instruction. Benchmarking four models shows rankings shift across preference types, exposing quality differences traditional evaluation obscures. We publicly release the dataset.

📄 PDF Abstract BibTeX arXiv:2606.29534

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

PrefSQA: Pairwise Preference Prediction for Speech Quality Assessment and the Critical Role of High Quality Datasets

2026-06-17 · Junyi Fan, Donald S. Williamson arxiv

Mean opinion scores (MOS) are widely used for speech quality assessment, yet scalar labels are sensitive to rater variability and listening test differences. This introduces labeling noise, which limits the reliability o…

Emotion-Aligned Generation in Diffusion Text to Speech Models via Preference-Guided Optimization

2025-09-29 · Jiacheng Shi, Hongfei Du, Yangfan He, Y. Alicia Hong 외 arxiv

Emotional text-to-speech seeks to convey affect while preserving intelligibility and prosody, yet existing methods rely on coarse labels or proxy classifiers and receive only utterance-level feedback. We introduce Emotio…

Text to Speech

Decoding the Ear: A Framework for Objectifying Expressiveness from Human Preference Through Efficient Alignment

2025-10-23 · Zhiyu Lin, Jingwen Yang, Jiale Zhao, Meng Liu 외 arxiv

Recent speech-to-speech (S2S) models generate intelligible speech but still lack natural expressiveness, largely due to the absence of a reliable evaluation metric. Existing approaches, such as subjective MOS ratings, lo…

Emotion Recognition

From Scores to Preferences: Redefining MOS Benchmarking for Speech Quality Reward Modeling

2025-10-01 · Yifei Cao, Changhao Jiang, Jiabao Zhuang, Jiajun Sun 외 arxiv

Assessing the perceptual quality of synthetic speech is crucial for guiding the development and refinement of speech generation models. However, it has traditionally relied on human subjective ratings such as the Mean Op…

AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation

2025-07-17 · Potsawee Manakul, Woody Haosheng Gan, Michael J. Ryan, Ali Sartaz Khan 외 arxiv

Current speech evaluation suffers from two critical limitations: the need and difficulty of designing specialized systems targeting individual audio characteristics, and poor correlation between automatic evaluation meth…

Speaker IdentificationPrompt Engineering