paper-with-me

Papers

Unsupervised Composable Representations for Audio

2024-08-19 · Giovanni Bindi, Philippe Esling

Current generative models are able to generate high-quality artefacts but have been shown to struggle with compositional reasoning, which can be defined as the ability to generate complex structures from simpler elements. In this paper, we focus on the problem of compositional representation learning for music data, specifically targeting the fully-unsupervised setting. We propose a simple and extensible framework that leverages an explicit compositional inductive bias, defined by a flexible auto-encoding objective that can leverage any of the current state-of-art generative models. We demonstrate that our framework, used with diffusion models, naturally addresses the task of unsupervised audio source separation, showing that our model is able to perform high-quality separation. Our findings reveal that our proposal achieves comparable or superior performance with respect to other blind source separation methods and, furthermore, it even surpasses current state-of-art supervised baselines on signal-to-interference ratio metrics. Additionally, by learning an a-posteriori masking diffusion model in the space of composable representations, we achieve a system capable of seamlessly performing unsupervised source separation, unconditional generation, and variation generation. Finally, as our proposal works in the latent space of pre-trained neural audio codecs, it also provides a lower computational cost with respect to other neural baselines.

📄 PDF Abstract BibTeX arXiv:2408.09792

Code (1)

ismir-24-sub/unsupervised_compositional_representations 공식 구현 pytorch

Tasks

Audio Source Separationblind source separationInductive BiasRepresentation Learning

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
Focus 설명 없음

Similar Papers 제목 키워드 기반

Unsupervised Improvement of Audio-Text Cross-Modal Representations

2023-05-03 · Zhepei Wang, Cem Subakan, Krishna Subramani, Junkai Wu 외

Recent advances in using language models to obtain cross-modal audio-text representations have overcome the limitations of conventional training approaches that use predefined labels. This has allowed the community to ma…

Acoustic Scene ClassificationClassificationScene Classificationzero-shot-classification+1

Speed Co-Augmentation for Unsupervised Audio-Visual Pre-training

2023-09-25 · Jiangliu Wang, Jianbo Jiao, Yibing Song, Stephen James 외

This work aims to improve unsupervised audio-visual pre-training. Inspired by the efficacy of data augmentation in visual contrastive learning, we propose a novel speed co-augmentation method that randomly changes the pl…

Contrastive LearningData AugmentationDiversity

Audio Word2Vec: Unsupervised Learning of Audio Segment Representations using Sequence-to-sequence Autoencoder

2016-03-03 · Yu-An Chung, Chao-Chung Wu, Chia-Hao Shen, Hung-Yi Lee 외

The vector representations of fixed dimensionality for words (in text) offered by Word2Vec have been shown to be very useful in many application scenarios, in particular due to the semantic information they carry. This p…

DecoderDenoisingDynamic Time Warping

Any-to-Any Generation via Composable Diffusion

2023-05-19 · NeurIPS 2023 11 · Zineng Tang, ZiYi Yang, Chenguang Zhu, Michael Zeng 외

We present Composable Diffusion (CoDi), a novel generative model capable of generating any combination of output modalities, such as language, image, video, or audio, from any combination of input modalities. Unlike exis…

Audio Generation

Unsupervised Modality-Transferable Video Highlight Detection with Representation Activation Sequence Learning

2024-03-14 · Tingtian Li, Zixun Sun, Xinyu Xiao

Identifying highlight moments of raw video materials is crucial for improving the efficiency of editing videos that are pervasive on internet platforms. However, the extensive work of manually labeling footage has create…

Contrastive LearningHighlight Detection