Dialog-context aware end-to-end speech recognition
Existing speech recognition systems are typically built at the sentence level, although it is known that dialog context, e.g. higher-level knowledge that spans across sentences or speakers, can help the processing of long conversations. The recent progress in end-to-end speech recognition systems promises to integrate all available information (e.g. acoustic, language resources) into a single model, which is then jointly optimized. It seems natural that such dialog context information should thus also be integrated into the end-to-end models to improve further recognition accuracy. In this work, we present a dialog-context aware speech recognition model, which explicitly uses context information beyond sentence-level information, in an end-to-end fashion. Our dialog-context model captures a history of sentence-level context so that the whole system can be trained with dialog-context information in an end-to-end manner. We evaluate our proposed approach on the Switchboard conversational speech corpus and show that our system outperforms a comparable sentence-level end-to-end speech recognition system.
Code (0)
등록된 구현이 없습니다.
Tasks
Sentencespeech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
Enhancing Dialogue Speech Recognition with Robust Contextual Awareness via Noise Representation Learning
Recent dialogue systems rely on turn-based spoken interactions, requiring accurate Automatic Speech Recognition (ASR). Errors in ASR can significantly impact downstream dialogue tasks. To address this, using dialogue con…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderRepresentation Learning+2Context-Aware Dialog Re-Ranking for Task-Oriented Dialog Systems
Dialog response ranking is used to rank response candidates by considering their relation to the dialog history. Although researchers have addressed this concept for open-domain dialogs, little attention has been focused…
Re-Rankingspeech-recognitionSpeech RecognitionEmotion-Aware Speech Generation with Character-Specific Voices for Comics
This paper presents an end-to-end pipeline for generating character-specific, emotion-aware speech from comics. The proposed system takes full comic volumes as input and produces speech aligned with each character's dial…
An Interaction-aware Attention Network for Speech Emotion Recognition in Spoken Dialogs
In this work, we propose an interaction-aware attention network (IAAN) that incorporate contextual information in the learned vocal representation through a novel attention mechanism. Our proposed method achieves 66.3% a…
Emotion RecognitionSpeech Emotion RecognitionReading the Mood Behind Words: Integrating Prosody-Derived Emotional Context into Socially Responsive VR Agents
In VR interactions with embodied conversational agents, users' emotional intent is often conveyed more by how something is said than by what is said. However, most VR agent pipelines rely on speech-to-text processing, di…
Speech Emotion Recognition