paper-with-me

홈 › Papers

Temporal Order Preserved Optimal Transport-based Cross-modal Knowledge Transfer Learning for ASR

2024-09-03 · Xugang Lu, Peng Shen, Yu Tsao, Hisashi Kawai

Transferring linguistic knowledge from a pretrained language model (PLM) to an acoustic model has been shown to greatly improve the performance of automatic speech recognition (ASR). However, due to the heterogeneous feature distributions in cross-modalities, designing an effective model for feature alignment and knowledge transfer between linguistic and acoustic sequences remains a challenging task. Optimal transport (OT), which efficiently measures probability distribution discrepancies, holds great potential for aligning and transferring knowledge between acoustic and linguistic modalities. Nonetheless, the original OT treats acoustic and linguistic feature sequences as two unordered sets in alignment and neglects temporal order information during OT coupling estimation. Consequently, a time-consuming pretraining stage is required to learn a good alignment between the acoustic and linguistic representations. In this paper, we propose a Temporal Order Preserved OT (TOT)-based Cross-modal Alignment and Knowledge Transfer (CAKT) (TOT-CAKT) for ASR. In the TOT-CAKT, local neighboring frames of acoustic sequences are smoothly mapped to neighboring regions of linguistic sequences, preserving their temporal order relationship in feature alignment and matching. With the TOT-CAKT model framework, we conduct Mandarin ASR experiments with a pretrained Chinese PLM for linguistic knowledge transfer. Our results demonstrate that the proposed TOT-CAKT significantly improves ASR performance compared to several state-of-the-art models employing linguistic knowledge transfer, and addresses the weaknesses of the original OT-based method in sequential feature alignment for ASR.

📄 PDF Abstract BibTeX arXiv:2409.02239

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)cross-modal alignmentspeech-recognitionSpeech RecognitionTransfer Learning

Similar Papers 제목 키워드 기반

BOTM: Echocardiography Segmentation via Bi-directional Optimal Token Matching

2025-05-23 · Zhihua Liu, Lei Tong, Xilin He, Che Liu 외

Existed echocardiography segmentation methods often suffer from anatomical inconsistency challenge caused by shape variation, partial observation and region ambiguity with similar intensity across 2D echocardiographic se…

AnatomySegmentation

Robot Policy Learning with Temporal Optimal Transport Reward

2024-10-29 · Yuwei Fu, Haichao Zhang, Di wu, Wei Xu 외

Reward specification is one of the most tricky problems in Reinforcement Learning, which usually requires tedious hand engineering in practice. One promising approach to tackle this challenge is to adopt existing expert …

Order-Preserving Wasserstein Distance for Sequence Matching

2017-07-01 · CVPR 2017 7 · Bing Su, Gang Hua

We present a new distance measure between sequences that can tackle local temporal distortion and periodic sequences with arbitrary starting points. Through viewing the instances of sequences as empirical samples of an u…

Cross-user activity recognition via temporal relation optimal transport

2024-03-12 · Xiaozhou Ye, Kevin I-Kai Wang

Current research on human activity recognition (HAR) mainly assumes that training and testing data are drawn from the same distribution to achieve a generalised model, which means all the data are considered to be indepe…

Activity RecognitionDomain AdaptationHuman Activity RecognitionRelation+2

MultistageOT: Multistage optimal transport infers trajectories from a snapshot of single-cell data

2025-02-07 · Magnus Tronstad, Johan Karlsson, Joakim S. Dahlin

Single-cell RNA-sequencing captures a temporal slice, or a snapshot, of a cell differentiation process. A major bioinformatical challenge is the inference of differentiation trajectories from a single snapshot, and metho…