paper-with-me

Papers

The Far Side of Failure: Investigating the Impact of Speech Recognition Errors on Subsequent Dementia Classification

2022-11-11 · Changye Li, Trevor Cohen, Serguei Pakhomov

Linguistic anomalies detectable in spontaneous speech have shown promise for various clinical applications including screening for dementia and other forms of cognitive impairment. The feasibility of deploying automated tools that can classify language samples obtained from speech in large-scale clinical settings depends on the ability to capture and automatically transcribe the speech for subsequent analysis. However, the impressive performance of self-supervised learning (SSL) automatic speech recognition (ASR) models with curated speech data is not apparent with challenging speech samples from clinical settings. One of the key questions for successfully applying ASR models for clinical applications is whether imperfect transcripts they generate provide sufficient information for downstream tasks to operate at an acceptable level of accuracy. In this study, we examine the relationship between the errors produced by several deep learning ASR systems and their impact on the downstream task of dementia classification. One of our key findings is that, paradoxically, ASR systems with relatively high error rates can produce transcripts that result in better downstream classification accuracy than classification based on verbatim transcripts.

📄 PDF Abstract BibTeX arXiv:2211.07430

Code (1)

linguisticanomalies/paradox-asr 공식 구현 pytorch

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)ClassificationSelf-Supervised Learningspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Investigating the Impact of ASR Errors on Spoken Implicit Discourse Relation Recognition

2022-10-01 · TU (COLING) 2022 10 · Linh The Nguyen, Dat Quoc Nguyen

We present an empirical study investigating the influence of automatic speech recognition (ASR) errors on the spoken implicit discourse relation recognition (IDRR) task. We construct a spoken dataset for this task based …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Relationspeech-recognition+1

Back Transcription as a Method for Evaluating Robustness of Natural Language Understanding Models to Speech Recognition Errors

2023-10-25 · Marek Kubis, Paweł Skórzewski, Marcin Sowański, Tomasz Ziętkiewicz

In a spoken dialogue system, an NLU model is preceded by a speech recognition system that can deteriorate the performance of natural language understanding. This paper proposes a method for investigating the impact of sp…

en-US domain classificationen-US Intent Classificationen-US Slot FillingNatural Language Understanding+4

Investigating the Lombard Effect Influence on End-to-End Audio-Visual Speech Recognition

2019-06-05 · Pingchuan Ma, Stavros Petridis, Maja Pantic

Several audio-visual speech recognition models have been recently proposed which aim to improve the robustness over audio-only models in the presence of noise. However, almost all of them ignore the impact of the Lombard…

Audio-Visual Speech Recognitionspeech-recognitionSpeech RecognitionVisual Speech Recognition

Investigating the Impact of Word Informativeness on Speech Emotion Recognition

2025-06-02 · Sofoklis Kakouros

In emotion recognition from speech, a key challenge lies in identifying speech signal segments that carry the most relevant acoustic variations for discerning specific emotions. Traditional approaches compute functionals…

Emotion RecognitionInformativenessLanguage ModelingLanguage Modelling+1

Investigating Cross-Domain Losses for Speech Enhancement

2020-10-20 · Sherif Abdulatif, Karim Armanious, Jayasankar T. Sajeev, Karim Guirguis 외

Recent years have seen a surge in the number of available frameworks for speech enhancement (SE) and recognition. Whether model-based or constructed via deep learning, these frameworks often rely in isolation on either t…

Deep LearningSpeech Enhancement