paper-with-me

홈 › Papers

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs

2026-07-07 · Albert Zeyer, Ralf Schlüter, Hermann Ney arxiv

Speech-to-text alignment means finding the temporal boundaries of each word in the audio. Some models provide such an alignment directly and others do not. Connectionist temporal classification (CTC) and transducer models have an alignment by construction, whereas attention-based encoder-decoders (AED) and speech large language models (LLMs) do not, and their word timings are usually read off the attention weights instead. All of these signals live on the encoder frame grid, which bounds their temporal precision. We study a generic gradient-based alignment that applies to any differentiable ASR model. We take the gradient of each teacher-forced token log probability with respect to the input, reduce it to a per-frame saliency, and decode the resulting matrix into word boundaries with a single dynamic-programming pass. The method needs no training, no model modification and no alignment heads, works across all model families including the speech LLMs, and aligns on the input grid rather than on the coarser encoder grid. We evaluate it on sixteen models from four families, on read (TIMIT) and spontaneous (Buckeye) speech, each against the model's own native or attention-based alignment. We find that the gradient yields a usable alignment for every model, that it is usually somewhat behind a strong native aligner but better where the native alignment is weak, as for the streaming models, and that its main disadvantage is the cost of one backward pass per token.

📄 PDF Abstract BibTeX arXiv:2607.06831

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Adaptive Inner Speech-Text Alignment for LLM-based Speech Translation

2025-03-13 · Henglyu Liu, Andong Chen, Kehai Chen, Xuefeng Bai 외

Recent advancement of large language models (LLMs) has led to significant breakthroughs across various tasks, laying the foundation for the development of LLM-based speech translation systems. Existing methods primarily …

Cross-Modal RetrievalTranslation

Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving

2025-05-24 · Jingran Xie, Xiang Li, Hui Wang, Yue Yu 외

Large language models (LLMs) have shown remarkable generalization across tasks, leading to increased interest in integrating speech with LLMs. These speech LLMs (SLLMs) typically use supervised fine-tuning to align speec…

Decoder

BLSP-KD: Bootstrapping Language-Speech Pre-training via Knowledge Distillation

2024-05-29 · Chen Wang, Minpeng Liao, Zhongqiang Huang, Jiajun Zhang

Recent end-to-end approaches have shown promise in extending large language models (LLMs) to speech inputs, but face limitations in directly assessing and optimizing alignment quality and fail to achieve fine-grained ali…

Instruction FollowingKnowledge Distillation

Efficient Training for Cross-lingual Speech Language Models

2026-04-13 · Yan Zhou, Qingkai Fang, Yun Hong, Yang Feng arxiv

Currently, large language models (LLMs) predominantly focus on the text modality. To enable more natural human-AI interaction, speech LLMs are emerging, but building effective end-to-end speech LLMs remains challenging d…

BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing

2023-09-02 · Chen Wang, Minpeng Liao, Zhongqiang Huang, Jinliang Lu 외

The emergence of large language models (LLMs) has sparked significant interest in extending their remarkable language capabilities to speech. However, modality alignment between speech and text still remains an open prob…

speech-recognitionSpeech RecognitionSpoken Language Understanding