paper-with-me

홈 › Papers

Acoustic-to-Word Models with Conversational Context Information

2019-05-21 · NAACL 2019 6 · Suyoun Kim, Florian Metze

Conversational context information, higher-level knowledge that spans across sentences, can help to recognize a long conversation. However, existing speech recognition models are typically built at a sentence level, and thus it may not capture important conversational context information. The recent progress in end-to-end speech recognition enables integrating context with other available information (e.g., acoustic, linguistic resources) and directly recognizing words from speech. In this work, we present a direct acoustic-to-word, end-to-end speech recognition model capable of utilizing the conversational context to better process long conversations. We evaluate our proposed approach on the Switchboard conversational speech corpus and show that our system outperforms a standard end-to-end speech recognition system.

📄 PDF Abstract BibTeX arXiv:1905.08796

Code (0)

등록된 구현이 없습니다.

Tasks

Sentencespeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

M2-CTTS: End-to-End Multi-scale Multi-modal Conversational Text-to-Speech Synthesis

2023-05-03 · Jinlong Xue, Yayue Deng, Fengping Wang, Ya Li 외

Conversational text-to-speech (TTS) aims to synthesize speech with proper prosody of reply based on the historical conversation. However, it is still a challenge to comprehensively model the conversation, and a majority …

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Improving OOV Detection and Resolution with External Language Models in Acoustic-to-Word ASR

2019-09-22 · Hirofumi Inaguma, Masato Mimura, Shinsuke Sakai, Tatsuya Kawahara

Acoustic-to-word (A2W) end-to-end automatic speech recognition (ASR) systems have attracted attention because of an extremely simplified architecture and fast decoding. To alleviate data sparseness issues due to infreque…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Gated Embeddings in End-to-End Speech Recognition for Conversational-Context Fusion

2019-06-27 · ACL 2019 7 · Suyoun Kim, Siddharth Dalmia, Florian Metze

We present a novel conversational-context aware end-to-end speech recognizer based on a gated neural network that incorporates conversational-context/word/speech embeddings. Unlike conventional speech recognition models,…

SentenceSentence Embeddingsspeech-recognitionSpeech Recognition

Super-Human Performance in Online Low-latency Recognition of Conversational Speech

2020-10-07 · Thai-Son Nguyen, Sebastian Stueker, Alex Waibel

Achieving super-human performance in recognizing human speech has been a goal for several decades, as researchers have worked on increasingly challenging tasks. In the 1990's it was discovered, that conversational speech…

Decoder

The Microsoft 2017 Conversational Speech Recognition System

2017-08-21 · W. Xiong, L. Wu, F. Alleva, J. Droppo 외

We describe the 2017 version of Microsoft's conversational speech recognition system, in which we update our 2016 system with recent developments in neural-network-based acoustic and language modeling to further advance …

Language ModelingLanguage Modellingspeech-recognitionSpeech Recognition