paper-with-me

Papers

Enabling Auditory Large Language Models for Automatic Speech Quality Evaluation

2024-09-25 · Siyin Wang, Wenyi Yu, Yudong Yang, Changli Tang, Yixuan Li, Jimin Zhuang, Xianzhao Chen, Xiaohai Tian, Jun Zhang, Guangzhi Sun, Lu Lu, Yuxuan Wang, Chao Zhang

Speech quality assessment typically requires evaluating audio from multiple aspects, such as mean opinion score (MOS) and speaker similarity (SIM) \etc., which can be challenging to cover using one small model designed for a single task. In this paper, we propose leveraging recently introduced auditory large language models (LLMs) for automatic speech quality assessment. By employing task-specific prompts, auditory LLMs are finetuned to predict MOS, SIM and A/B testing results, which are commonly used for evaluating text-to-speech systems. Additionally, the finetuned auditory LLM is able to generate natural language descriptions assessing aspects like noisiness, distortion, discontinuity, and overall quality, providing more interpretable outputs. Extensive experiments have been performed on the NISQA, BVCC, SOMOS and VoxSim speech quality datasets, using open-source auditory LLMs such as SALMONN, Qwen-Audio, and Qwen2-Audio. For the natural language descriptions task, a commercial model Google Gemini 1.5 Pro is also evaluated. The results demonstrate that auditory LLMs achieve competitive performance compared to state-of-the-art task-specific small models in predicting MOS and SIM, while also delivering promising results in A/B testing and natural language descriptions. Our data processing scripts and finetuned model checkpoints can be found at https://github.com/bytedance/SALMONN.

📄 PDF Abstract BibTeX arXiv:2409.16644

Code (1)

bytedance/salmonn 공식 구현 pytorch

Tasks

text-to-speechText to Speech

Similar Papers 제목 키워드 기반

Auditory-Based Data Augmentation for End-to-End Automatic Speech Recognition

2022-04-08 · Zehai Tu, Jack Deadman, Ning Ma, Jon Barker

End-to-end models have achieved significant improvement on automatic speech recognition. One common method to improve performance of these models is expanding the data-space through data augmentation. Meanwhile, human au…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentationspeech-recognition+1

SALMONN: Towards Generic Hearing Abilities for Large Language Models

2023-10-20 · Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen 외

Hearing is arguably an essential ability of artificial intelligence (AI) agents in the physical world, which refers to the perception and understanding of general auditory information consisting of at least three types o…

Audio captioningAutomatic Speech RecognitionEmotion RecognitionLanguage Modelling+8

Scaling Auditory Cognition via Test-Time Compute in Audio Language Models

2025-03-30 · Ting Dang, Yan Gao, Hong Jia

Large language models (LLMs) have shown exceptional versatility in natural language processing, prompting recent efforts to extend their multimodal capabilities to speech processing through the development of audio large…

speech-recognitionSpeech Recognition

From Sounds to Scenes: A Benchmark for Evaluating Context-Aware Auditory Scene Understanding in Large Audio Language Models

2026-06-24 · Pengfei Zhang, Hoang H Nguyen, Kazi Shaharair Sharif, Yutong Song 외 arxiv

Recent Large Audio Language Models (LALMs) have achieved remarkable progress in audio perceptual tasks across individual acoustic layers, including speech, sound, and music. However, existing benchmarks predominantly eva…

Scene UnderstandingQuestion Answering

DeSTA2.5-Audio: Toward General-Purpose Large Audio Language Model with Self-Generated Cross-Modal Alignment

2025-07-03 · Ke-Han Lu, Zhehuai Chen, Szu-Wei Fu, Chao-Han Huck Yang 외

We introduce DeSTA2.5-Audio, a general-purpose Large Audio Language Model (LALM) designed for robust auditory perception and instruction-following, without requiring task-specific audio instruction-tuning. Recent LALMs t…

cross-modal alignmentInstruction FollowingLanguage ModelingLanguage Modelling+1