paper-with-me

Papers

tinyCLAP: Distilling Constrastive Language-Audio Pretrained Models

2023-11-24 · Francesco Paissan, Elisabetta Farella

Contrastive Language-Audio Pretraining (CLAP) became of crucial importance in the field of audio and speech processing. Its employment ranges from sound event detection to text-to-audio generation. However, one of the main limitations is the considerable amount of data required in the training process and the overall computational complexity during inference. This paper investigates how we can reduce the complexity of contrastive language-audio pre-trained models, yielding an efficient model that we call tinyCLAP. We derive an unimodal distillation loss from first principles and explore how the dimensionality of the shared, multimodal latent space can be reduced via pruning. TinyCLAP uses only 6% of the original Microsoft CLAP parameters with a minimal reduction (less than 5%) in zero-shot classification performance across the three sound event detection datasets on which it was tested

📄 PDF Abstract BibTeX arXiv:2311.14517

Code (0)

등록된 구현이 없습니다.

Tasks

Audio GenerationEvent DetectionSound Event Detectionzero-shot-classificationZero-Shot Learning

Similar Papers 제목 키워드 기반

Score Distillation Sampling for Audio: Source Separation, Synthesis, and Beyond

2025-05-07 · Jessie Richter-Powell, Antonio Torralba, Jonathan Lorraine

We introduce Audio-SDS, a generalization of Score Distillation Sampling (SDS) to text-conditioned audio diffusion models. While SDS was initially designed for text-to-3D generation using image diffusion, its core idea of…

3D GenerationAudio Source SeparationText to 3D

XDBERT: Distilling Visual Information to BERT from Cross-Modal Systems to Improve Language Understanding

2022-04-15 · ACL 2022 5 · Chan-Jan Hsu, Hung-Yi Lee, Yu Tsao

Transformer-based models are widely used in natural language understanding (NLU) tasks, and multimodal transformers have been effective in visual-language tasks. This study explores distilling visual information from pre…

Natural Language Understanding

CLIPSonic: Text-to-Audio Synthesis with Unlabeled Videos and Pretrained Language-Vision Models

2023-06-16 · Hao-Wen Dong, Xiaoyu Liu, Jordi Pons, Gautam Bhattacharya 외

Recent work has studied text-to-audio synthesis using large amounts of paired text-audio data. However, audio recordings with high-quality text annotations can be difficult to acquire. In this work, we approach text-to-a…

Audio Synthesis

CoLLAT: On Adding Fine-grained Audio Understanding to Language Models using Token-Level Locked-Language Tuning

2023-09-21 · NeurIPS 2023 11

Humans can easily understand various audio concepts, but conventional audio classification models fail due to their inability to predict unseen classes during training. To address this challenge, recent literature has ex…

Few-Shot Pidgin Text Adaptation via Contrastive Fine-Tuning

2022-10-01 · COLING 2022 10 · Ernie Chang, Jesujoba O. Alabi, David Ifeoluwa Adelani, Vera Demberg

The surging demand for multilingual dialogue systems often requires a costly labeling process for each language addition. For low resource languages, human annotators are continuously tasked with the adaptation of resour…

Text Generation