paper-with-me

홈 › Papers

Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs

2025-12-18 · Sara Papi, Javier Garcia Gilabert, Zachary Hopton, Vilém Zouhar, Carlos Escolano, Gerard I. Gállego, Jorge Iranzo-Sánchez, Ahrii Kim, Dominik Macháček, Patricia Schmidtova, Maike Züfle arxiv

As Large Language Models (LLMs) expand beyond text, integrating speech as a native modality has given rise to SpeechLLMs, which directly process spoken language and enable speech-to-text translation (ST) and other downstream tasks, bypassing traditional transcription-based pipelines. Whether this integration improves ST quality over established cascaded architectures, however, remains an open question. We present Hearing to Translate, the first comprehensive test suite rigorously benchmarking 6 state-of-the-art SpeechLLMs against 16 strong direct and cascade systems that couple leading speech foundation models (SFM), with multilingual LLMs. Our analysis spans 16 benchmarks, 13 language pairs, and 9 challenging conditions, including disfluent, noisy, and long-form speech. Across this extensive evaluation, we find that cascaded systems remain the most reliable solution overall, but most recent SpeechLLMs can match or even outperform cascades in various settings while SFMs lag behind both, highlighting that integrating an LLM, either within the model or in a pipeline, is essential for high-quality speech translation.

📄 PDF Abstract BibTeX arXiv:2512.16378

Code (0)

등록된 구현이 없습니다.

Tasks

Speech-to-Text Translation

Similar Papers 제목 키워드 기반

From TOWER to SPIRE: Adding the Speech Modality to a Text-Only LLM

2025-03-13 · Kshitij Ambilduke, Ben Peters, Sonal Sannigrahi, Anil Keshwani 외

Large language models (LLMs) have shown remarkable performance and generalization capabilities across multiple languages and tasks, making them very attractive targets for multi-modality integration (e.g., images or spee…

Translation

Real-Time Sign Language Gestures to Speech Transcription using Deep Learning

2025-08-18 · Brandone Fonya, Clarence Worrell arxiv

Communication barriers pose significant challenges for individuals with hearing and speech impairments, often limiting their ability to effectively interact in everyday environments. This project introduces a real-time a…

Text-To-Speech Synthesis

Advancing Hearing Assessment: An ASR-Based Frequency-Specific Speech Test for Diagnosing Presbycusis

2025-05-28 · Stefan Bleeck

Traditional audiometry often fails to fully characterize the functional impact of hearing loss on speech understanding, particularly supra-threshold deficits and frequency-specific perception challenges in conditions lik…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Diagnosticspeech-recognition+1

Artifact-free Sound Quality in DNN-based Closed-loop Systems for Audio Processing

2025-01-07 · Chuan Wen, Guy Torfs, Sarah Verhulst

Recent advances in deep neural networks (DNNs) have significantly improved various audio processing applications, including speech enhancement, synthesis, and hearing aid algorithms. DNN-based closed-loop systems have ga…

Speech Enhancement

Geometry-Constrained EEG Channel Selection for Brain-Assisted Speech Enhancement

2024-09-19 · Keying Zuo, Qingtian Xu, Jie Zhang, ZhenHua Ling

Brain-assisted speech enhancement (BASE) aims to extract the target speaker in complex multi-talker scenarios using electroencephalogram (EEG) signals as an assistive modality, as the auditory attention of the listener c…

channel selectionEEGElectroencephalogram (EEG)Speech Enhancement