paper-with-me

홈 › Papers

Neural domain alignment for spoken language recognition based on optimal transport

2023-10-20 · Xugang Lu, Peng Shen, Yu Tsao, Hisashi Kawai

Domain shift poses a significant challenge in cross-domain spoken language recognition (SLR) by reducing its effectiveness. Unsupervised domain adaptation (UDA) algorithms have been explored to address domain shifts in SLR without relying on class labels in the target domain. One successful UDA approach focuses on learning domain-invariant representations to align feature distributions between domains. However, disregarding the class structure during the learning process of domain-invariant representations can result in over-alignment, negatively impacting the classification task. To overcome this limitation, we propose an optimal transport (OT)-based UDA algorithm for a cross-domain SLR, leveraging the distribution geometry structure-aware property of OT. An OT-based discrepancy measure on a joint distribution over feature and label information is considered during domain alignment in OT-based UDA. Our previous study discovered that completely aligning the distributions between the source and target domains can introduce a negative transfer, where classes or irrelevant classes from the source domain map to a different class in the target domain during distribution alignment. This negative transfer degrades the performance of the adaptive model. To mitigate this issue, we introduce coupling-weighted partial optimal transport (POT) within our UDA framework for SLR, where soft weighting on the OT coupling based on transport cost is adaptively set during domain alignment. A cross-domain SLR task was used in the experiments to evaluate the proposed UDA. The results demonstrated that our proposed UDA algorithm significantly improved the performance over existing UDA algorithms in a cross-channel SLR task.

📄 PDF Abstract BibTeX arXiv:2310.13471

Code (0)

등록된 구현이 없습니다.

Tasks

Domain AdaptationUnsupervised Domain Adaptation

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…
SLR Please enter a description about the method here

Similar Papers 제목 키워드 기반

Partial Coupling of Optimal Transport for Spoken Language Identification

2022-03-31 · Xugang Lu, Peng Shen, Yu Tsao, Hisashi Kawai

In order to reduce domain discrepancy to improve the performance of cross-domain spoken language identification (SLID) system, as an unsupervised domain adaptation (UDA) method, we have proposed a joint distribution alig…

Domain AdaptationLanguage IdentificationSpoken language identificationUnsupervised Domain Adaptation

ALAS: Measuring Latent Speech-Text Alignment For Spoken Language Understanding In Multimodal LLMs

2025-05-26 · Pooneh Mousavi, Yingzhi Wang, Mirco Ravanelli, Cem Subakan

Large Language Models (LLMs) are widely used in Spoken Language Understanding (SLU). Recent SLU models process audio directly by adapting speech input into LLMs for better multimodal learning. A key consideration for the…

cross-modal alignmentEmotion RecognitionQuestion AnsweringSpoken Language Understanding

Optimal Transport Regularization for Speech Text Alignment in Spoken Language Models

2025-08-11 · Wenze Xu, Chun Wang, Jiazhen Yu, Sheng Chen 외 arxiv

Spoken Language Models (SLMs), which extend Large Language Models (LLMs) to perceive speech inputs, have gained increasing attention for their potential to advance speech understanding tasks. However, despite recent prog…

ADIMA: Abuse Detection In Multilingual Audio

2022-02-16 · Vikram Gupta, Rini Sharon, Ramit Sawhney, Debdoot Mukherjee

Abusive content detection in spoken text can be addressed by performing Automatic Speech Recognition (ASR) and leveraging advancements in natural language processing. However, ASR models introduce latency and often perfo…

Abuse DetectionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognition+1

Gujarati-English Code-Switching Speech Recognition using ensemble prediction of spoken language

2024-03-12 · Yash Sharma, Basil Abraham, Preethi Jyothi

An important and difficult task in code-switched speech recognition is to recognize the language, as lots of words in two languages can sound similar, especially in some accents. We focus on improving performance of end-…

Automatic Speech Recognitionspeech-recognitionSpeech Recognition