The importance of fillers for text representations of speech transcripts
While being an essential component of spoken language, fillers (e.g."um" or "uh") often remain overlooked in Spoken Language Understanding (SLU) tasks. We explore the possibility of representing them with deep contextualised embeddings, showing improvements on modelling spoken language and two downstream tasks - predicting a speaker's stance and expressed confidence.
Code (0)
등록된 구현이 없습니다.
Tasks
Spoken Language UnderstandingSimilar Papers 제목 키워드 기반
Evaluating Sampling-based Filler Insertion with Spontaneous TTS
Inserting fillers (such as “um”, “like”) to clean speech text has a rich history of study. One major application is to make dialogue systems sound more spontaneous. The ambiguity of filler occurrence and inter-speaker di…
Importance-related Fillers Improve the Classification Accuracy of the Response Time Concealed Information Test in a Crime Scenario
Purpose. The Response Time Concealed Information Test (RT-CIT) can reveal when a person recognizes a relevant item among other irrelevant items, based on comparatively slower responding. Therefore, if a person is conceal…
DiagnosticOpen-Ended Question AnsweringComedicSpeech: Text To Speech For Stand-up Comedies in Low-Resource Scenarios
Text to Speech (TTS) models can generate natural and high-quality speech, but it is not expressive enough when synthesizing speech with dramatic expressiveness, such as stand-up comedies. Considering comedians have diver…
Rhythmtext-to-speechText to SpeechMind the Pause: Disfluency-Aware Objective Tuning for Multilingual Speech Correction with LLMs
Automatic Speech Recognition (ASR) transcripts often contain disfluencies, such as fillers, repetitions, and false starts, which reduce readability and hinder downstream applications like chatbots and voice assistants. I…
Contrastive LearningSpeech RecognitionData AugmentationDo We Still Need Automatic Speech Recognition for Spoken Language Understanding?
Spoken language understanding (SLU) tasks are usually solved by first transcribing an utterance with automatic speech recognition (ASR) and then feeding the output to a text-based model. Recent advances in self-supervise…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationnamed-entity-recognition+7