paper-with-me

홈 › Papers

BAT: Better Audio Transformer Guided by Convex Gated Probing

2026-02-18 · Houtan Ghaffari, Lukas Rauch, Christoph Scholz, Paul Devos arxiv

Probing is widely adopted in computer vision to faithfully evaluate self-supervised learning (SSL) embeddings, as finetuning may misrepresent their inherent quality. In contrast, audio SSL models still rely on finetuning because simple probing fails to unlock their full potential and alters their rankings when competing on AudioSet. Hence, a robust and efficient probing mechanism is required to guide the trajectory of audio SSL towards reliable and reproducible methods. We introduce Convex Gated Probing (CGP), a prototype-based method that significantly closes the gap between finetuning and probing in audio. CGP efficiently utilizes all frozen layers via a gating mechanism and exposes the location of latent task-relevant information. Guided by CGP as a reliable post-hoc evaluation probe, we rework the entire SSL pipeline of current best performing audio models that use legacy implementations of prior SSL methods. By refining data preprocessing, model architecture, and pretraining recipe, we introduce Better Audio Transformer (BAT), and establish new SOTA on audio benchmarks.

📄 PDF Abstract BibTeX arXiv:2602.16305

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised Learning

Similar Papers 제목 키워드 기반

Learning Audio-guided Video Representation with Gated Attention for Video-Text Retrieval

2025-04-03 · CVPR 2025 1 · Boseung Jeong, Jicheol Park, Sungyeon Kim, Suha Kwak

Video-text retrieval, the task of retrieving videos based on a textual query or vice versa, is of paramount importance for video understanding and multimodal information retrieval. Recent methods in this area rely primar…

Information RetrievalRepresentation LearningRetrievalText Retrieval+2

AudioMAE++: learning better masked audio representations with SwiGLU FFNs

2025-07-14 · Sarthak Yadav, Sergios Theodoridis, Zheng-Hua Tan arxiv

Masked Autoencoders (MAEs) trained on audio spectrogram patches have emerged as a prominent approach for learning self-supervised audio representations. While several recent papers have evaluated key aspects of training …

Audio Classification

Convexity-based Pruning of Speech Representation Models

2024-08-16 · Teresa Dorszewski, Lenka Tětková, Lars Kai Hansen

Speech representation models based on the transformer architecture and trained by self-supervised learning have shown great promise for solving tasks such as speech and speaker recognition, keyword spotting, emotion dete…

Keyword SpottingSelf-Supervised LearningSpeaker Recognition

Audio Mamba: Selective State Spaces for Self-Supervised Audio Representations

2024-06-04 · Sarthak Yadav, Zheng-Hua Tan

Despite its widespread adoption as the prominent neural architecture, the Transformer has spurred several independent lines of work to address its limitations. One such approach is selective state space models, which hav…

Language ModellingMambaState Space Models

Sounding Video Generator: A Unified Framework for Text-guided Sounding Video Generation

2023-03-29 · Jiawei Liu, Weining Wang, Sihan Chen, Xinxin Zhu 외

As a combination of visual and audio signals, video is inherently multi-modal. However, existing video generation methods are primarily intended for the synthesis of visual frames, whereas audio signals in realistic vide…

Audio GenerationContrastive LearningDecoderVideo Generation