paper-with-me

홈 › Papers

Token-Weighted RNN-T for Learning from Flawed Data

2024-06-26 · Gil Keren, Wei Zhou, Ozlem Kalinli

ASR models are commonly trained with the cross-entropy criterion to increase the probability of a target token sequence. While optimizing the probability of all tokens in the target sequence is sensible, one may want to de-emphasize tokens that reflect transcription errors. In this work, we propose a novel token-weighted RNN-T criterion that augments the RNN-T objective with token-specific weights. The new objective is used for mitigating accuracy loss from transcriptions errors in the training data, which naturally appear in two settings: pseudo-labeling and human annotation errors. Experiments results show that using our method for semi-supervised learning with pseudo-labels leads to a consistent accuracy improvement, up to 38% relative. We also analyze the accuracy degradation resulting from different levels of WER in the reference transcription, and show that token-weighted RNN-T is suitable for overcoming this degradation, recovering 64%-99% of the accuracy loss.

📄 PDF Abstract BibTeX arXiv:2406.18108

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

GRAIL: Gradient-Reweighted Advantages for Reinforcement Learning with Verifiable Rewards

2026-06-03 · Tej Deep Pala, Vernon Toh, Soujanya Poria arxiv

Reinforcement learning with verifiable rewards (e.g. GRPO) is now a common way to improve mathematical reasoning in Large Language Models (LLMs). However, current methods usually broadcast one sequence-level advantage to…

Reinforcement LearningMathematical Reasoning

FAPO: Flawed-Aware Policy Optimization for Efficient and Reliable Reasoning

2025-10-26 · Yuyang Ding, Chi Zhang, Juntao Li, Haibin Lin 외 arxiv

Reinforcement learning with verifiable rewards (RLVR) has emerged as a promising paradigm for enhancing the reasoning capabilities of large language models (LLMs). In this context, models explore reasoning trajectories a…

Reinforcement Learning

TAME: Token Attribution and Masking for Emergent misalignment

2026-09-15 · Md Rayhanul Masud, Md Rizwan Parvez arxiv

Fine-tuning an aligned language model on narrow, flawed data can induce harmful behavior far outside the training domain, known as emergent misalignment (EM). Prior work has localized EM in model weights, activations, an…

Large Reasoning Models Learn Better Alignment from Flawed Thinking

2025-10-01 · ShengYun Peng, Eric Smith, Ivan Evtimov, Song Jiang 외 arxiv

Large reasoning models (LRMs) "think" by generating structured chain-of-thought (CoT) before producing a final answer, yet they still lack the ability to reason critically about safety alignment and are easily biased whe…

Reinforcement Learning

Negative Self-Distillation: Learning to Reason by Avoiding Flaws

2026-09-10 · Rongcan Pei, Zhepei Wei, Shuyao Xu, Xinyu Zhu 외 hf

On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for large language model (LLM) self-improvement, allowing models to act as their own teachers by leveraging privileged information such as ground-truth…

Reinforcement Learning