paper-with-me

홈 › Papers

Expectation and Locality Effects in the Prediction of Disfluent Fillers and Repairs in English Speech

2019-06-01 · NAACL 2019 6 · Samvit Dammalapati, Rajakrishnan Rajkumar, Sumeet Agarwal

This study examines the role of three influential theories of language processing, \textit{viz.}, Surprisal Theory, Uniform Information Density (UID) hypothesis and Dependency Locality Theory (DLT), in predicting disfluencies in speech production. To this end, we incorporate features based on lexical surprisal, word duration and DLT integration and storage costs into logistic regression classifiers aimed to predict disfluencies in the Switchboard corpus of English conversational speech. We find that disfluencies occur in the face of upcoming difficulties and speakers tend to handle this by lessening cognitive load before disfluencies occur. Further, we see that reparandums behave differently from disfluent fillers possibly due to the lessening of the cognitive load also happening in the word choice of the reparandum, i.e., in the disfluency itself. While the UID hypothesis does not seem to play a significant role in disfluency prediction, lexical surprisal and DLT costs do give promising results in explaining language production. Further, we also find that as a means to lessen cognitive load for upcoming difficulties speakers take more time on words preceding disfluencies, making duration a key element in understanding disfluencies.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Logistic Regression Logistic Regression, despite its name, is a linear model for classification rather than regression. Logistic regression is also known in the literature as logit regression,…

Similar Papers 제목 키워드 기반

Mind the Pause: Disfluency-Aware Objective Tuning for Multilingual Speech Correction with LLMs

2026-05-12 · Deepak Kumar, Baban Gain, Asif Ekbal arxiv

Automatic Speech Recognition (ASR) transcripts often contain disfluencies, such as fillers, repetitions, and false starts, which reduce readability and hinder downstream applications like chatbots and voice assistants. I…

Contrastive LearningSpeech RecognitionData Augmentation

Noisy-context surprisal as a human sentence processing cost model

2017-04-01 · EACL 2017 4 · Richard Futrell, Roger Levy

We use the noisy-channel theory of human sentence comprehension to develop an incremental processing cost model that unifies and extends key features of expectation-based and memory-based models. In this model, which we …

Sentence

Planning and Generating Natural and Diverse Disfluent Texts as Augmentation for Disfluency Detection

2020-11-01 · EMNLP 2020 11 · Jingfeng Yang, Diyi Yang, Zhaoran Ma

Existing approaches to disfluency detection heavily depend on human-annotated data. Numbers of data augmentation methods have been proposed to alleviate the dependence on labeled data. However, current augmentation appro…

Data Augmentation

DISCO: A Large Scale Human Annotated Corpus for Disfluency Correction in Indo-European Languages

2023-10-25 · Vineet Bhat, Preethi Jyothi, Pushpak Bhattacharyya

Disfluency correction (DC) is the process of removing disfluent elements like fillers, repetitions and corrections from spoken utterances to create readable and interpretable text. DC is a vital post-processing step appl…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognition+1

What makes a good pause? Investigating the turn-holding effects of fillers

2023-05-03 · Bing'er Jiang, Erik Ekstedt, Gabriel Skantze

Filled pauses (or fillers), such as "uh" and "um", are frequent in spontaneous speech and can serve as a turn-holding cue for the listener, indicating that the current speaker is not done yet. In this paper, we use the r…

Position