paper-with-me

홈 › Papers

Careless Whisper: Speech-to-Text Hallucination Harms

2024-02-12 · Allison Koenecke, Anna Seo Gyeong Choi, Katelyn X. Mei, Hilke Schellmann, Mona Sloane

Speech-to-text services aim to transcribe input audio as accurately as possible. They increasingly play a role in everyday life, for example in personal voice assistants or in customer-company interactions. We evaluate Open AI's Whisper, a state-of-the-art automated speech recognition service outperforming industry competitors, as of 2023. While many of Whisper's transcriptions were highly accurate, we find that roughly 1\% of audio transcriptions contained entire hallucinated phrases or sentences which did not exist in any form in the underlying audio. We thematically analyze the Whisper-hallucinated content, finding that 38\% of hallucinations include explicit harms such as perpetuating violence, making up inaccurate associations, or implying false authority. We then study why hallucinations occur by observing the disparities in hallucination rates between speakers with aphasia (who have a lowered ability to express themselves using speech and voice) and a control group. We find that hallucinations disproportionately occur for individuals who speak with longer shares of non-vocal durations -- a common symptom of aphasia. We call on industry practitioners to ameliorate these language-model-based hallucinations in Whisper, and to raise awareness of potential biases amplified by hallucinations in downstream applications of speech-to-text models.

📄 PDF Abstract BibTeX arXiv:2402.08021

Code (1)

koenecke/hallucination_harms 공식 구현

Tasks

HallucinationLanguage ModelingLanguage Modellingspeech-recognitionSpeech RecognitionSpeech-to-Text

Methods 이 논문이 사용한 방법론

Golden Queue Managers 설명 없음

Similar Papers 제목 키워드 기반

Calm-Whisper: Reduce Whisper Hallucination On Non-Speech By Calming Crazy Heads Down

2025-05-19 · Yingzhi Wang, Anas Alhmoud, Saad Alsahly, Muhammad Alqurishi 외

OpenAI's Whisper has achieved significant success in Automatic Speech Recognition. However, it has consistently been found to exhibit hallucination issues, particularly in non-speech segments, which limits its broader ap…

Automatic Speech RecognitionDecoderHallucinationspeech-recognition+1

Investigation of Whisper ASR Hallucinations Induced by Non-Speech Audio

2025-01-20 · Mateusz Barański, Jan Jasiński, Julitta Bartolewska, Stanisław Kacprzak 외

Hallucinations of deep neural models are amongst key challenges in automatic speech recognition (ASR). In this paper, we investigate hallucinations of the Whisper ASR model induced by non-speech audio segments present du…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Whisper Hallucination Detection and Mitigation via Hidden Representation Steering and Sparse AutoEncoders

2026-06-05 · Georgii Aparin, Vadim Popov, Tasnima Sadekova, Assel Yermekova arxiv

Whisper, a widely adopted ASR model, is known to suffer from hallucinations - coherent transcriptions generated for non-speech audio entirely disconnected from the input. We investigate whether hallucinations can be dete…

Reducing Hallucinated Transcripts in Whisper via Hallucination Space Projection

2026-09-03 · Maryam Abbasihafshejani, Murtuza Jadliwala arxiv

Whisper is a widely used foundation model for automatic speech recognition (ASR), but its generative decoder can produce fluent hallucinated transcripts for inputs containing little or no speech. We propose a training-fr…

Speech Recognition

From Text Metrics to Model Internals: A Study of Whisper ASR Hallucination Detection

2026-06-22 · Jan Jasiński, Mateusz Barański, Julitta Bartolewska, Marcin Witkowski 외 arxiv

Hallucinations of ASR models - fluent transcriptions with no basis in audio - degrade system performance and pose risks in downstream applications. Robust detection of such errors remains a challenge. This paper studies …