paper-with-me

Papers

Investigating Design Choices in Joint-Embedding Predictive Architectures for General Audio Representation Learning

2024-05-14 · Alain Riou, Stefan Lattner, Gaëtan Hadjeres, Geoffroy Peeters

This paper addresses the problem of self-supervised general-purpose audio representation learning. We explore the use of Joint-Embedding Predictive Architectures (JEPA) for this task, which consists of splitting an input mel-spectrogram into two parts (context and target), computing neural representations for each, and training the neural network to predict the target representations from the context representations. We investigate several design choices within this framework and study their influence through extensive experiments by evaluating our models on various audio classification benchmarks, including environmental sounds, speech and music downstream tasks. We focus notably on which part of the input data is used as context or target and show experimentally that it significantly impacts the model's quality. In particular, we notice that some effective design choices in the image domain lead to poor performance on audio, thus highlighting major differences between these two modalities.

📄 PDF Abstract BibTeX arXiv:2405.08679

Code (1)

SonyCSLParis/audio-representations 공식 구현 pytorch

Tasks

Audio ClassificationRepresentation Learning

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

CrossJEPA: Cross-Modal Joint-Embedding Predictive Architecture for Efficient 3D Representation Learning from 2D Images

2025-11-23 · Avishka Perera, Kumal Hewagamage, Saeedha Nazar, Kavishka Abeywardana 외 arxiv

Image-to-point cross-modal learning has emerged to address the scarcity of large-scale 3D datasets in 3D representation learning. However, current methods that leverage 2D data often result in large, slow-to-train models…

Self-Supervised LearningRepresentation LearningKnowledge DistillationPoint Clouds

Tiny, always-on and fragile: Bias propagation through design choices in on-device machine learning workflows

2022-01-19 · Wiebke Toussaint, Aaron Yi Ding, Fahim Kawsar, Akhil Mathur

Billions of distributed, heterogeneous and resource constrained IoT devices deploy on-device machine learning (ML) for private, fast and offline inference on personal data. On-device ML is highly context dependent, and s…

Keyword Spotting

LiDAR: Sensing Linear Probing Performance in Joint Embedding SSL Architectures

2023-12-07 · Vimal Thilak, Chen Huang, Omid Saremi, Laurent Dinh 외

Joint embedding (JE) architectures have emerged as a promising avenue for acquiring transferable data representations. A key obstacle to using JE methods, however, is the inherent challenge of evaluating learned represen…

Imitating Human Behaviour with Diffusion Models

2023-01-25 · Tim Pearce, Tabish Rashid, Anssi Kanervisto, Dave Bignell 외

Diffusion models have emerged as powerful generative models in the text-to-image domain. This paper studies their application as observation-to-action models for imitating human behaviour in sequential environments. Huma…

Analysis of Twitter Users' Lifestyle Choices using Joint Embedding Model

2021-04-07 · Tunazzina Islam, Dan Goldwasser

Multiview representation learning of data can help construct coherent and contextualized users' representations on social media. This paper suggests a joint embedding model, incorporating users' social and textual inform…

Representation Learning