paper-with-me

홈 › Papers

Building Robust Spoken Language Understanding by Cross Attention between Phoneme Sequence and ASR Hypothesis

2022-03-22 · Zexun Wang, Yuquan Le, Yi Zhu, Yuming Zhao, Mingchao Feng, Meng Chen, Xiaodong He

Building Spoken Language Understanding (SLU) robust to Automatic Speech Recognition (ASR) errors is an essential issue for various voice-enabled virtual assistants. Considering that most ASR errors are caused by phonetic confusion between similar-sounding expressions, intuitively, leveraging the phoneme sequence of speech can complement ASR hypothesis and enhance the robustness of SLU. This paper proposes a novel model with Cross Attention for SLU (denoted as CASLU). The cross attention block is devised to catch the fine-grained interactions between phoneme and word embeddings in order to make the joint representations catch the phonetic and semantic features of input simultaneously and for overcoming the ASR errors in downstream natural language understanding (NLU) tasks. Extensive experiments are conducted on three datasets, showing the effectiveness and competitiveness of our approach. Additionally, We also validate the universality of CASLU and prove its complementarity when combining with other robust SLU techniques.

📄 PDF Abstract BibTeX arXiv:2203.12067

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Natural Language Understandingspeech-recognitionSpeech RecognitionSpoken Language UnderstandingWord Embeddings

Similar Papers 제목 키워드 기반

Transferring SLU Models in Novel Domains

2019-05-01 · ICLR 2019 5 · Yaohua Tang, Kaixiang Mo, Qian Xu, Chao Zhang 외

Spoken language understanding (SLU) is a critical component in building dialogue systems. When building models for novel natural language domains, a major challenge is the lack of data in the new domains, no matter wheth…

Intent RecognitionMeta-Learningslot-fillingSlot Filling+2

Finstreder: Simple and fast Spoken Language Understanding with Finite State Transducers using modern Speech-to-Text models

2022-06-29 · Daniel Bermuth, Alexander Poeppel, Wolfgang Reif

In Spoken Language Understanding (SLU) the task is to extract important information from audio commands, like the intent of what a user wants the system to do and special entities like locations or numbers. This paper pr…

Intent ClassificationSlot FillingSpeech-to-TextSpoken Language Understanding

A Preliminary Evaluation of ChatGPT for Zero-shot Dialogue Understanding

2023-04-09 · Wenbo Pan, Qiguang Chen, Xiao Xu, Wanxiang Che 외

Zero-shot dialogue understanding aims to enable dialogue to track the user's needs without any training data, which has gained increasing attention. In this work, we investigate the understanding ability of ChatGPT for z…

Dialogue State TrackingDialogue Understandingslot-fillingSlot Filling+1

On Building Spoken Language Understanding Systems for Low Resourced Languages

2022-05-25 · NAACL (SIGMORPHON) 2022 7 · Akshat Gupta

Spoken dialog systems are slowly becoming and integral part of the human experience due to their various advantages over textual interfaces. Spoken language understanding (SLU) systems are fundamental building blocks of …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)intent-classificationIntent Classification+3

Improving End-to-End Models for Set Prediction in Spoken Language Understanding

2022-01-28 · Hong-Kwang J. Kuo, Zoltan Tuske, Samuel Thomas, Brian Kingsbury 외

The goal of spoken language understanding (SLU) systems is to determine the meaning of the input speech signal, unlike speech recognition which aims to produce verbatim transcripts. Advances in end-to-end (E2E) speech mo…

Data AugmentationDecoderspeech-recognitionSpeech Recognition+1