paper-with-me

Papers

End-to-End Spoken Language Understanding for Generalized Voice Assistants

2021-06-16 · Michael Saxon, Samridhi Choudhary, Joseph P. McKenna, Athanasios Mouchtaris

End-to-end (E2E) spoken language understanding (SLU) systems predict utterance semantics directly from speech using a single model. Previous work in this area has focused on targeted tasks in fixed domains, where the output semantic structure is assumed a priori and the input speech is of limited complexity. In this work we present our approach to developing an E2E model for generalized SLU in commercial voice assistants (VAs). We propose a fully differentiable, transformer-based, hierarchical system that can be pretrained at both the ASR and NLU levels. This is then fine-tuned on both transcription and semantic classification losses to handle a diverse set of intent and argument combinations. This leads to an SLU system that achieves significant improvements over baselines on a complex internal generalized VA dataset with a 43% improvement in accuracy, while still meeting the 99% accuracy benchmark on the popular Fluent Speech Commands dataset. We further evaluate our model on a hard test set, exclusively containing slot arguments unseen in training, and demonstrate a nearly 20% improvement, showing the efficacy of our approach in truly demanding VA scenarios.

📄 PDF Abstract BibTeX arXiv:2106.09009

Code (0)

등록된 구현이 없습니다.

Tasks

Spoken Language Understanding

Similar Papers 제목 키워드 기반

MultiVox: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions

2025-07-14 · Ramaneswaran Selvakumar, Ashish Seth, Nishit Anand, Utkarsh Tyagi 외 arxiv

The rapid progress of Large Language Models (LLMs) has empowered omni models to act as voice assistants capable of understanding spoken dialogues. These models can process multimodal inputs beyond text, such as speech an…

Sonos Voice Control Bias Assessment Dataset: A Methodology for Demographic Bias Assessment in Voice Assistants

2024-05-14 · Chloé Sekkat, Fanny Leroy, Salima Mdhaffar, Blake Perry Smith 외

Recent works demonstrate that voice assistants do not perform equally well for everyone, but research on demographic robustness of speech technologies is still scarce. This is mainly due to the rarity of large datasets w…

Automatic Speech RecognitionDiversityspeech-recognitionSpeech Recognition+1

VoiceAgentBench: Are Voice Assistants ready for agentic tasks?

2025-10-09 · Dhruv Jain, Harshit Shukla, Gautam Rajeev, Ashish Kulkarni 외 arxiv

Large scale Speech Language Models have enabled voice assistants capable of understanding natural spoken queries and performing complex tasks. However, existing speech benchmarks largely focus on isolated capabilities su…

Adversarial RobustnessQuestion AnsweringVoice Conversion

Noise Robust Named Entity Understanding for Voice Assistants

2020-05-29 · NAACL 2021 4 · Deepak Muralidharan, Joel Ruben Antony Moniz, Sida Gao, Xiao Yang 외

Named Entity Recognition (NER) and Entity Linking (EL) play an essential role in voice assistant interaction, but are challenging due to the special difficulties associated with spoken user queries. In this paper, we pro…

domain classificationEntity Linkingnamed-entity-recognitionNamed Entity Recognition+5

VoiceBench: Benchmarking LLM-Based Voice Assistants

2024-10-22 · Yiming Chen, Xianghu Yue, Chen Zhang, Xiaoxue Gao 외

Building on the success of large language models (LLMs), recent advancements such as GPT-4o have enabled real-time speech interactions through LLM-based voice assistants, offering a significantly improved user experience…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)BenchmarkingGeneral Knowledge+2