Personal VAD 2.0: Optimizing Personal Voice Activity Detection for On-Device Speech Recognition
Personalization of on-device speech recognition (ASR) has seen explosive growth in recent years, largely due to the increasing popularity of personal assistant features on mobile devices and smart home speakers. In this work, we present Personal VAD 2.0, a personalized voice activity detector that detects the voice activity of a target speaker, as part of a streaming on-device ASR system. Although previous proof-of-concept studies have validated the effectiveness of Personal VAD, there are still several critical challenges to address before this model can be used in production: first, the quality must be satisfactory in both enrollment and enrollment-less scenarios; second, it should operate in a streaming fashion; and finally, the model size should be small enough to fit a limited latency and CPU/Memory budget. To meet the multi-faceted requirements, we propose a series of novel designs: 1) advanced speaker embedding modulation methods; 2) a new training paradigm to generalize to enrollment-less conditions; 3) architecture and runtime optimizations for latency and resource restrictions. Extensive experiments on a realistic speech recognition system demonstrated the state-of-the-art performance of our proposed method.
Code (0)
등록된 구현이 없습니다.
Tasks
Action DetectionActivity DetectionCPUspeech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
Comparative Analysis of Personalized Voice Activity Detection Systems: Assessing Real-World Effectiveness
Voice activity detection (VAD) is a critical component in various applications such as speech recognition, speech enhancement, and hands-free communication systems. With the increasing demand for personalized and context…
Action DetectionActivity DetectionSpeech Enhancementspeech-recognition+1Self-supervised Pretraining for Robust Personalized Voice Activity Detection in Adverse Conditions
In this paper, we propose the use of self-supervised pretraining on a large unlabelled data set to improve the performance of a personalized voice activity detection (VAD) model in adverse conditions. We pretrain a long …
Action DetectionActivity DetectionDenoisingHyWA: Hypernetwork Weight Adapting Personalized Voice Activity Detection
Personalized Voice Activity Detection (PVAD) systems activate only in response to a specific target speaker. Speaker-conditioning methods are employed to inject information about the target speaker into a VAD pipeline, t…
Activity DetectionBetter Together: Dialogue Separation and Voice Activity Detection for Audio Personalization in TV
In TV services, dialogue level personalization is key to meeting user preferences and needs. When dialogue and background sounds are not separately available from the production stage, Dialogue Separation (DS) can estima…
Action DetectionActivity DetectionArray Configuration-Agnostic Personal Voice Activity Detection Based on Spatial Coherence
Personal voice activity detection has received increased attention due to the growing popularity of personal mobile devices and smart speakers. PVAD is often an integral element to speech enhancement and recognition for …
Action DetectionActivity DetectionSpeech Enhancement