paper-with-me

홈 › Papers

Analyzing the Importance of Blank for CTC-Based Knowledge Distillation

2025-06-02 · Benedikt Hilmes, Nick Rossenbach, Ralf Schlüter

With the rise of large pre-trained foundation models for automatic speech recognition new challenges appear. While the performance of these models is good, runtime and cost of inference increases. One approach to make use of their strength while retaining efficiency is to distill their knowledge to smaller models during training. In this work, we explore different CTC-based distillation variants, focusing on blank token handling. We show that common approaches like blank elimination do not always work off the shelf. We explore new blank selection patterns as a potential sweet spot between standard knowledge distillation and blank elimination mechanisms. Through the introduction of a symmetric selection method, we are able to remove the CTC loss during knowledge distillation with minimal to no performance degradation. With this, we make the training independent from target labels, potentially allowing for distillation on untranscribed audio data.

📄 PDF Abstract BibTeX arXiv:2506.01503

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionKnowledge Distillationspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

Knowledge Distillation A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions.…
CTC Loss 설명 없음

Similar Papers 제목 키워드 기반

An Empirical Analysis of Likelihood-Weighting Simulation on a Large, Multiply-Connected Belief Network

2013-03-27 · Michael Shwe, Gregory F. Cooper

We analyzed the convergence properties of likelihood- weighting algorithms on a two-level, multiply connected, belief-network representation of the QMR knowledge base of internal medicine. Specifically, on two difficult …

Diagnostic

CTC Blank Triggered Dynamic Layer-Skipping for Efficient CTC-based Speech Recognition

2024-01-04 · JunFeng Hou, Peiyao Wang, Jincheng Zhang, Meng Yang 외

Deploying end-to-end speech recognition models with limited computing resources remains challenging, despite their impressive performance. Given the gradual increase in model size and the wide range of model applications…

Knowledge Distillationspeech-recognitionSpeech Recognition

Accelerating Evolution Through Gene Masking and Distributed Search

2023-02-13 · Hormoz Shahrzad, Risto Miikkulainen

In building practical applications of evolutionary computation (EC), two optimizations are essential. First, the parameters of the search method need to be tuned to the domain in order to balance exploration and exploita…

Cards Against AI: Predicting Humor in a Fill-in-the-blank Party Game

2022-10-24 · Dan Ofer, Dafna Shahaf

Humor is an inherently social phenomenon, with humorous utterances shaped by what is socially and culturally accepted. Understanding humor is an important NLP challenge, with many applications to human-computer interacti…

Feature Importance

Towards Accurate Cross-Domain In-Bed Human Pose Estimation

2021-10-07 · Mohamed Afham, Udith Haputhanthri, Jathurshan Pradeepkumar, Mithunjha Anandakumar 외

Human behavioral monitoring during sleep is essential for various medical applications. Majority of the contactless human pose estimation algorithms are based on RGB modality, causing ineffectiveness in in-bed pose estim…

Data AugmentationKnowledge DistillationPose Estimation