paper-with-me

Papers

Bridging the Gap between Audio and Text using Parallel-attention for User-defined Keyword Spotting

2024-08-07 · Youkyum Kim, Jaemin Jung, Jihwan Park, Byeong-Yeol Kim, Joon Son Chung

This paper proposes a novel user-defined keyword spotting framework that accurately detects audio keywords based on text enrollment. Since audio data possesses additional acoustic information compared to text, there are discrepancies between these two modalities. To address this challenge, we present ParallelKWS, which utilises self- and cross-attention in a parallel architecture to effectively capture information both within and across the two modalities. We further propose a phoneme duration-based alignment loss that enforces the sequential correspondence between audio and text features. Extensive experimental results demonstrate that our proposed method achieves state-of-the-art performance on several benchmark datasets in both seen and unseen domains, without incorporating extra data beyond the dataset used in previous studies.

📄 PDF Abstract BibTeX arXiv:2408.03593

Code (0)

등록된 구현이 없습니다.

Tasks

Keyword Spotting

Similar Papers 제목 키워드 기반

cross-modal fusion techniques for utterance-level emotion recognition from text and speech

2023-02-05 · Jiachen Luo, Huy Phan, Joshua Reiss

Multimodal emotion recognition (MER) is a fundamental complex research problem due to the uncertainty of human emotional expression and the heterogeneity gap between different modalities. Audio and text modalities are pa…

Emotion RecognitionMultimodal Emotion Recognition

Diffusion-Link: Diffusion Probabilistic Model for Bridging the Audio-Text Modality Gap

2025-10-13 · KiHyun Nam, Jongmin Choi, Hyeongkeun Lee, Jungwoo Heo 외 arxiv

Contrastive audio-language pretraining yields powerful joint representations, yet a persistent audio-text modality gap limits the benefits of coupling multimodal encoders with large language models (LLMs). We present Dif…

Audio captioning

Connecting the Dots between Audio and Text without Parallel Data through Visual Knowledge Transfer

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Machines that can represent and describe environmental soundscapes have practical potential, e.g., for audio tagging and captioning. Prevailing learning paradigms of audio-text connections have been relying on parallel a…

Audio ClassificationAudio TaggingRetrievalTransfer Learning+1

Connecting the Dots between Audio and Text without Parallel Data through Visual Knowledge Transfer

2021-12-16 · NAACL 2022 7 · Yanpeng Zhao, Jack Hessel, Youngjae Yu, Ximing Lu 외

Machines that can represent and describe environmental soundscapes have practical potential, e.g., for audio tagging and captioning systems. Prevailing learning paradigms have been relying on parallel audio-text data, wh…

Audio ClassificationAudio TaggingRetrievalTransfer Learning+1

Q-TriM: Question-Guided Tri-Modal Attention for Audio-Visual Question Answering

2026-07-04 · SungHun Kim, SeungJun Baek arxiv

Audio-Visual Question Answering (AVQA) extends classical VQA by requiring joint reasoning over video and synchronized audio. However, many AVQA systems rely on deeply stacked layers of self- and cross attention across te…

Audio-visual Question Answering