Seq-2-Seq based Refinement of ASR Output for Spoken Name Capture
Person name capture from human speech is a difficult task in human-machine conversations. In this paper, we propose a novel approach to capture the person names from the caller utterances in response to the prompt "say and spell your first/last name". Inspired from work on spell correction, disfluency removal and text normalization, we propose a lightweight Seq-2-Seq system which generates a name spell from a varying user input. Our proposed method outperforms the strong baseline which is based on LM-driven rule-based approach.
Code (0)
등록된 구현이 없습니다.
Tasks
Text NormalizationSimilar Papers 제목 키워드 기반
Lost in Transcription: How Speech-to-Text Errors Derail Code Understanding
Code understanding is a foundational capability in software engineering tools and developer workflows. However, most existing systems are designed for English-speaking users interacting via keyboards, which limits access…
Speech RecognitionQuestion AnsweringMedical Spoken Named Entity Recognition
Spoken Named Entity Recognition (NER) aims to extract named entities from speech and categorise them into types like person, location, organization, etc. In this work, we present VietMed-NER - the first spoken NER datase…
named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+1MTL-SLT: Multi-Task Learning for Spoken Language Tasks
Language understanding in speech-based systems has attracted extensive interest from both academic and industrial communities in recent years with the growing demand for voice-based applications. Prior works focus on ind…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModellingMulti-Task Learning+4How Does That Sound? Multi-Language SpokenName2Vec Algorithm Using Speech Generation and Deep Learning
Searching for information about a specific person is an online activity frequently performed by many users. In most cases, users are aided by queries containing a name and sending back to the web search engines for findi…
English-Indonesian Neural Machine Translation for Spoken Language Domains
In this work, we conduct a study on Neural Machine Translation (NMT) for English-Indonesian (EN-ID) and Indonesian-English (ID-EN). We focus on spoken language domains, namely colloquial and speech languages. We build NM…
Domain AdaptationMachine TranslationNMTTranslation