paper-with-me

Papers

T-CLAP: Temporal-Enhanced Contrastive Language-Audio Pretraining

2024-04-27 · Yi Yuan, Zhuo Chen, Xubo Liu, Haohe Liu, Xuenan Xu, Dongya Jia, Yuanzhe Chen, Mark D. Plumbley, Wenwu Wang

Contrastive language-audio pretraining~(CLAP) has been developed to align the representations of audio and language, achieving remarkable performance in retrieval and classification tasks. However, current CLAP struggles to capture temporal information within audio and text features, presenting substantial limitations for tasks such as audio retrieval and generation. To address this gap, we introduce T-CLAP, a temporal-enhanced CLAP model. We use Large Language Models~(LLMs) and mixed-up strategies to generate temporal-contrastive captions for audio clips from extensive audio-text datasets. Subsequently, a new temporal-focused contrastive loss is designed to fine-tune the CLAP model by incorporating these synthetic data. We conduct comprehensive experiments and analysis in multiple downstream tasks. T-CLAP shows improved capability in capturing the temporal relationship of sound events and outperforms state-of-the-art models by a significant margin.

📄 PDF Abstract BibTeX arXiv:2404.17806

Code (0)

등록된 구현이 없습니다.

Tasks

Retrieval

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

GEmo-CLAP: Gender-Attribute-Enhanced Contrastive Language-Audio Pretraining for Accurate Speech Emotion Recognition

2023-06-13 · Yu Pan, Yanni Hu, Yuguang Yang, Wen Fei 외

Contrastive cross-modality pretraining has recently exhibited impressive success in diverse fields, whereas there is limited research on their merits in speech emotion recognition (SER). In this paper, we propose GEmo-CL…

AttributeContrastive LearningEmotion RecognitionMulti-Task Learning+2

Do Audio-Language Models Understand Linguistic Variations?

2024-10-21 · Ramaneswaran Selvakumar, Sonal Kumar, Hemant Kumar Giri, Nishit Anand 외

Open-vocabulary audio language models (ALMs), like Contrastive Language Audio Pretraining (CLAP), represent a promising new paradigm for audio-text retrieval using natural language queries. In this paper, for the first t…

Contrastive LearningNatural Language QueriesRetrievalText Retrieval+1

tinyCLAP: Distilling Constrastive Language-Audio Pretrained Models

2023-11-24 · Francesco Paissan, Elisabetta Farella

Contrastive Language-Audio Pretraining (CLAP) became of crucial importance in the field of audio and speech processing. Its employment ranges from sound event detection to text-to-audio generation. However, one of the ma…

Audio GenerationEvent DetectionSound Event Detectionzero-shot-classification+1

Enhancing Audio-Language Models through Self-Supervised Post-Training with Text-Audio Pairs

2024-08-17 · Anshuman Sinha, Camille Migozzi, Aubin Rey, Chao Zhang

Research on multi-modal contrastive learning strategies for audio and text has rapidly gained interest. Contrastively trained Audio-Language Models (ALMs), such as CLAP, which establish a unified representation across au…

Audio ClassificationContrastive LearningRetrievalZero-shot Audio Classification

Exploring the User Experience of AI-Assisted Sound Searching Systems for Creative Workflows

2025-04-22 · Haohe Liu, Thomas Deacon, Wenwu Wang, Matt Paradis 외

Locating the right sound effect efficiently is an important yet challenging topic for audio production. Most current sound-searching systems rely on pre-annotated audio labels created by humans, which can be time-consumi…

valid