Locality enhanced dynamic biasing and sampling strategies for contextual ASR
Automatic Speech Recognition (ASR) still face challenges when recognizing time-variant rare-phrases. Contextual biasing (CB) modules bias ASR model towards such contextually-relevant phrases. During training, a list of biasing phrases are selected from a large pool of phrases following a sampling strategy. In this work we firstly analyse different sampling strategies to provide insights into the training of CB for ASR with correlation plots between the bias embeddings among various training stages. Secondly, we introduce a neighbourhood attention (NA) that localizes self attention (SA) to the nearest neighbouring frames to further refine the CB output. The results show that this proposed approach provides on average a 25.84% relative WER improvement on LibriSpeech sets and rare-word evaluation compared to the baseline.
Code (0)
등록된 구현이 없습니다.
Tasks
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
Unbiasing Enhanced Sampling on a High-dimensional Free Energy Surface with Deep Generative Model
Biased enhanced sampling methods utilizing collective variables (CVs) are powerful tools for sampling conformational ensembles. Due to high intrinsic dimensions, efficiently generating conformational ensembles for comple…
Density EstimationAdaptive Sampling Methods for Molecular Dynamics in the Era of Machine Learning
Molecular Dynamics (MD) simulations are fundamental computational tools for the study of proteins and their free energy landscapes. However, sampling protein conformational changes through MD simulations is challenging d…
Tangent Space Least Adaptive Clustering
The biasing of dynamical simulations along collective variables uncovered by unsupervised learning has become a standard approach in analysis of molecular systems. However, despite parallels with reinforcement learning (…
Clusteringreinforcement-learningReinforcement LearningReinforcement Learning (RL)Towards Debiasing Temporal Sentence Grounding in Video
The temporal sentence grounding in video (TSGV) task is to locate a temporal moment from an untrimmed video, to match a language query, i.e., a sentence. Without considering bias in moment annotations (e.g., start and en…
SentenceTemporal Sentence GroundingReweighted Manifold Learning of Collective Variables from Enhanced Sampling Simulations
Enhanced sampling methods are indispensable in computational physics and chemistry, where atomistic simulations cannot exhaustively sample the high-dimensional configuration space of dynamical systems due to the sampling…