Analyzing the Importance of Blank for CTC-Based Knowledge Distillation
With the rise of large pre-trained foundation models for automatic speech recognition new challenges appear. While the performance of these models is good, runtime and cost of inference increases. One approach to make use of their strength while retaining efficiency is to distill their knowledge to smaller models during training. In this work, we explore different CTC-based distillation variants, focusing on blank token handling. We show that common approaches like blank elimination do not always work off the shelf. We explore new blank selection patterns as a potential sweet spot between standard knowledge distillation and blank elimination mechanisms. Through the introduction of a symmetric selection method, we are able to remove the CTC loss during knowledge distillation with minimal to no performance degradation. With this, we make the training independent from target labels, potentially allowing for distillation on untranscribed audio data.
Code (0)
등록된 구현이 없습니다.
Tasks
Automatic Speech RecognitionKnowledge Distillationspeech-recognitionSpeech RecognitionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
An Empirical Analysis of Likelihood-Weighting Simulation on a Large, Multiply-Connected Belief Network
We analyzed the convergence properties of likelihood- weighting algorithms on a two-level, multiply connected, belief-network representation of the QMR knowledge base of internal medicine. Specifically, on two difficult …
DiagnosticCTC Blank Triggered Dynamic Layer-Skipping for Efficient CTC-based Speech Recognition
Deploying end-to-end speech recognition models with limited computing resources remains challenging, despite their impressive performance. Given the gradual increase in model size and the wide range of model applications…
Knowledge Distillationspeech-recognitionSpeech RecognitionAccelerating Evolution Through Gene Masking and Distributed Search
In building practical applications of evolutionary computation (EC), two optimizations are essential. First, the parameters of the search method need to be tuned to the domain in order to balance exploration and exploita…
Cards Against AI: Predicting Humor in a Fill-in-the-blank Party Game
Humor is an inherently social phenomenon, with humorous utterances shaped by what is socially and culturally accepted. Understanding humor is an important NLP challenge, with many applications to human-computer interacti…
Feature ImportanceTowards Accurate Cross-Domain In-Bed Human Pose Estimation
Human behavioral monitoring during sleep is essential for various medical applications. Majority of the contactless human pose estimation algorithms are based on RGB modality, causing ineffectiveness in in-bed pose estim…
Data AugmentationKnowledge DistillationPose Estimation