paper-with-me

Papers

Pseudo2Real: Task Arithmetic for Pseudo-Label Correction in Automatic Speech Recognition

2025-10-09 · Yi-Cheng Lin, Yu-Hsuan Li Liang, Hsuan Su, Tzu-Quan Lin, Shang-Tse Chen, Yun-Nung Chen, Hung-yi Lee arxiv

Robust ASR under domain shift is crucial because real-world systems encounter unseen accents and domains with limited labeled data. Although pseudo-labeling offers a practical workaround, it often introduces systematic, accent-specific errors that filtering fails to fix. We ask: How can we correct these recurring biases without target ground truth? We propose a simple parameter-space correction: in a source domain containing both real and pseudo-labeled data, two ASR models are fine-tuned from the same initialization, one on ground-truth labels and the other on pseudo-labels, and their weight difference forms a correction vector that captures pseudo-label biases. When applied to a pseudo-labeled target model, this vector enhances recognition, achieving up to a 35% relative Word Error Rate (WER) reduction on AfriSpeech-200 across ten African accents with the Whisper tiny model.

📄 PDF Abstract BibTeX arXiv:2510.08047

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Recognition

Similar Papers 제목 키워드 기반

You can't handle the (dirty) truth: Data-centric insights improve pseudo-labeling

2024-06-19 · Nabeel Seedat, Nicolas Huynh, Fergus Imrie, Mihaela van der Schaar

Pseudo-labeling is a popular semi-supervised learning technique to leverage unlabeled data when labeled samples are scarce. The generation and selection of pseudo-labels heavily rely on labeled data. Existing approaches …

PseudoAugment: Learning to Use Unlabeled Data for Data Augmentation in Point Clouds

2022-10-24 · Zhaoqi Leng, Shuyang Cheng, Benjamin Caine, Weiyue Wang 외

Data augmentation is an important technique to improve data efficiency and save labeling cost for 3D detection in point clouds. Yet, existing augmentation policies have so far been designed to only utilize labeled data, …

Data AugmentationPseudo Labelvehicle detection

Expectation Maximization Pseudo Labels

2023-05-02 · MouCheng Xu, Yukun Zhou, Chen Jin, Marius de Groot 외

In this paper, we study pseudo-labelling. Pseudo-labelling employs raw inferences on unlabelled data as pseudo-labels for self-training. We elucidate the empirical successes of pseudo-labelling by establishing a link bet…

Segmentation

Pseudo-Loss Confidence Metric for Semi-Supervised Few-Shot Learning

2021-01-01 · ICCV 2021 10 · Kai Huang, Jie Geng, Wen Jiang, Xinyang Deng 외

Semi-supervised few-shot learning is developed to train a classifier that can adapt to new tasks with limited labeled data and a fixed quantity of unlabeled data. Most semi-supervised few-shot learning methods select…

Few-Shot Learning

Refined Pseudo labeling for Source-free Domain Adaptive Object Detection

2023-03-07 · Siqi Zhang, Lu Zhang, Zhiyong Liu

Domain adaptive object detection (DAOD) assumes that both labeled source data and unlabeled target data are available for training, but this assumption does not always hold in real-world scenarios. Thus, source-free DAOD…

object-detectionObject DetectionPseudo Label