paper-with-me

홈 › Papers

Joint Multimodal Contrastive Learning for Robust Spoken Term Detection and Keyword Spotting

2025-12-16 · Ramesh Gundluru, Shubham Gupta, Sri Rama Murty K arxiv

Acoustic Word Embeddings (AWEs) improve the efficiency of speech retrieval tasks such as Spoken Term Detection (STD) and Keyword Spotting (KWS). However, existing approaches suffer from limitations, including unimodal supervision, disjoint optimization of audio-audio and audio-text alignment, and the need for task-specific models. To address these shortcomings, we propose a joint multimodal contrastive learning framework that unifies both acoustic and cross-modal supervision in a shared embedding space. Our approach simultaneously optimizes: (i) audio-text contrastive learning, inspired by the CLAP loss, to align audio and text representations and (ii) audio-audio contrastive learning, via Deep Word Discrimination (DWD) loss, to enhance intra-class compactness and inter-class separation. The proposed method outperforms existing AWE baselines on word discrimination task while flexibly supporting both STD and KWS. To our knowledge, this is the first comprehensive approach of its kind.

📄 PDF Abstract BibTeX arXiv:2512.14115

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningKeyword Spotting

Similar Papers 제목 키워드 기반

Towards Spoken Language Understanding via Multi-level Multi-grained Contrastive Learning

2024-05-31 · Xuxin Cheng, Wanshi Xu, Zhihong Zhu, Hongxiang Li 외

Spoken language understanding (SLU) is a core task in task-oriented dialogue systems, which aims at understanding the user's current goal through constructing semantic frames. SLU usually consists of two subtasks, includ…

Contrastive LearningIntent Detectionslot-fillingSlot Filling+2

Joint Multiple Intent Detection and Slot Filling with Supervised Contrastive Learning and Self-Distillation

2023-08-28 · Nguyen Anh Tu, Hoang Thi Thu Uyen, Tu Minh Phuong, Ngo Xuan Bach

Multiple intent detection and slot filling are two fundamental and crucial tasks in spoken language understanding. Motivated by the fact that the two tasks are closely related, joint models that can detect intents and ex…

Contrastive LearningIntent DetectionSemantic Frame Parsingslot-filling+2

BDetCLIP: Multimodal Prompting Contrastive Test-Time Backdoor Detection

2024-05-24 · Yuwei Niu, Shuo He, Qi Wei, Zongyu Wu 외

Multimodal contrastive learning methods (e.g., CLIP) have shown impressive zero-shot classification performance due to their strong ability to joint representation learning for visual and textual modalities. However, rec…

Contrastive LearningLanguage ModellingRepresentation Learningzero-shot-classification+1

Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face Conversation

2024-06-12 · Se Jin Park, Chae Won Kim, Hyeongseop Rha, Minsu Kim 외

In this paper, we introduce a novel Face-to-Face spoken dialogue model. It processes audio-visual speech from user input and generates audio-visual speech as the response, marking the initial step towards creating an ava…

ChatbotLanguage ModelingLanguage ModellingLarge Language Model

TextMI: Textualize Multimodal Information for Integrating Non-verbal Cues in Pre-trained Language Models

2023-03-27 · Md Kamrul Hasan, Md Saiful Islam, Sangwu Lee, Wasifur Rahman 외

Pre-trained large language models have recently achieved ground-breaking performance in a wide variety of language understanding tasks. However, the same model can not be applied to multimodal behavior understanding task…

Humor DetectionMultimodal Sentiment AnalysisSarcasm DetectionSentiment Analysis