paper-with-me

Papers

Listening or Reading? Evaluating Speech Awareness in Chain-of-Thought Speech-to-Text Translation

2025-10-03 · Jacobo Romero-Díaz, Gerard I. Gállego, Oriol Pareras, Federico Costa, Javier Hernando, Cristina España-Bonet arxiv

Speech-to-Text Translation (S2TT) systems built from Automatic Speech Recognition (ASR) and Text-to-Text Translation (T2TT) modules face two major limitations: error propagation and the inability to exploit prosodic or other acoustic cues. Chain-of-Thought (CoT) prompting has recently been introduced, with the expectation that jointly accessing speech and transcription will overcome these issues. Analyzing CoT through attribution methods, robustness evaluations with corrupted transcripts, and prosody-awareness, we find that it largely mirrors cascaded behavior, relying mainly on transcripts while barely leveraging speech. Simple training interventions, such as adding Direct S2TT data or noisy transcript injection, enhance robustness and increase speech attribution. These findings challenge the assumed advantages of CoT and highlight the need for architectures that explicitly integrate acoustic information into translation.

📄 PDF Abstract BibTeX arXiv:2510.03115

Code (0)

등록된 구현이 없습니다.

Tasks

Speech-to-Text TranslationSpeech Recognition

Similar Papers 제목 키워드 기반

Spoken SQuAD: A Study of Mitigating the Impact of Speech Recognition Errors on Listening Comprehension

2018-04-01 · Chia-Hsuan Li, Szu-Lin Wu, Chi-Liang Liu, Hung-Yi Lee

Reading comprehension has been widely studied. One of the most representative reading comprehension tasks is Stanford Question Answering Dataset (SQuAD), on which machine is already comparable with human. On the other ha…

Question AnsweringReading Comprehensionspeech-recognitionSpeech Recognition+1

Listening while Speaking and Visualizing: Improving ASR through Multimodal Chain

2019-06-03 · Johanes Effendi, Andros Tjandra, Sakriani Sakti, Satoshi Nakamura

Previously, a machine speech chain, which is based on sequence-to-sequence deep learning, was proposed to mimic speech perception and production behavior. Such chains separately processed listening and speaking by automa…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationImage Captioning+7

Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?

2024-08-12 · Roshan Sharma, Suwon Shon, Mark Lindsey, Hira Dhamyal 외

Reference summaries for abstractive speech summarization require human annotation, which can be performed by listening to an audio recording or by reading textual transcripts of the recording. In this paper, we examine w…

Retrieval

Towards Optimizing OCR for Accessibility

2022-06-21 · Peya Mowar, Tanuja Ganu, Saikat Guha

Visual cues such as structure, emphasis, and icons play an important role in efficient information foraging by sighted individuals and make for a pleasurable reading experience. Blind, low-vision and other print-disabled…

Optical Character Recognition (OCR)text-to-speechText to Speech

Speech language models lack important brain-relevant semantics

2023-11-08 · Subba Reddy Oota, Emin Çelik, Fatma Deniz, Mariya Toneva

Despite known differences between reading and listening in the brain, recent work has shown that text-based language models predict both text-evoked and speech-evoked brain activity to an impressive degree. This poses th…

Language ModelingLanguage Modelling