paper-with-me

Papers

MetaAudio: A Few-Shot Audio Classification Benchmark

2022-04-05 · Calum Heggan, Sam Budgett, Timothy Hospedales, Mehrdad Yaghoobi

Currently available benchmarks for few-shot learning (machine learning with few training examples) are limited in the domains they cover, primarily focusing on image classification. This work aims to alleviate this reliance on image-based benchmarks by offering the first comprehensive, public and fully reproducible audio based alternative, covering a variety of sound domains and experimental settings. We compare the few-shot classification performance of a variety of techniques on seven audio datasets (spanning environmental sounds to human-speech). Extending this, we carry out in-depth analyses of joint training (where all datasets are used during training) and cross-dataset adaptation protocols, establishing the possibility of a generalised audio few-shot classification algorithm. Our experimentation shows gradient-based meta-learning methods such as MAML and Meta-Curvature consistently outperform both metric and baseline methods. We also demonstrate that the joint training routine helps overall generalisation for the environmental sound databases included, as well as being a somewhat-effective method of tackling the cross-dataset/domain setting.

📄 PDF Abstract BibTeX arXiv:2204.02121

Code (1)

cheggan/metaaudio-a-few-shot-audio-classification-benchmark 공식 구현 pytorch

Tasks

Audio ClassificationClassificationFew-Shot Audio ClassificationFew-Shot Learningimage-classificationImage ClassificationMeta-Learning

Methods 이 논문이 사용한 방법론

MAML 설명 없음

Similar Papers 제목 키워드 기반

Prototypical Contrastive Learning For Improved Few-Shot Audio Classification

2025-09-12 · Christos Sgouropoulos, Christos Nikou, Stefanos Vlachos, Vasileios Theiou 외 arxiv

Few-shot learning has emerged as a powerful paradigm for training models with limited labeled data, addressing challenges in scenarios where large-scale annotation is impractical. While extensive research has been conduc…

Few-Shot Audio ClassificationContrastive LearningFew-Shot Learning

SpurAudio: A Benchmark for Studying Shortcut Learning in Few-Shot Audio Classification

2026-05-13 · Giries Abu Ayoub, Morad Tukan, Loay Mualem arxiv

Few-shot classification (FSC) is widely used for learning from limited labeled data, yet most evaluations implicitly assume that target concepts are independent of contextual cues. In real-world settings, however, exampl…

Few-Shot Image ClassificationFew-Shot Audio Classification

Text-to-feature diffusion for audio-visual few-shot learning

2023-09-07 · Otniel-Bogdan Mercea, Thomas Hummel, A. Sophia Koepke, Zeynep Akata

Training deep learning models for video classification from audio-visual data commonly requires immense amounts of labeled training data collected via a costly process. A challenging and underexplored, yet much cheaper, …

ClassificationFew-Shot LearningVideo Classification

MVEB: Massive Video Embedding Benchmark

2026-06-12 · Adnan El Assadi, Roman Solomatin, Isaac Chung, Chenghao Xiao 외 arxiv

We introduce the Massive Video Embedding Benchmark (MVEB), a 23-task benchmark for video embeddings spanning classification, zero-shot classification, clustering, pair classification, retrieval, and video-centric questio…

Question Answering

CALM: Class-Conditional Sparse Attention Vectors for Large Audio-Language Models

2026-02-06 · Videet Mehta, Liming Wang, Hilde Kuehne, Rogerio Feris 외 arxiv

Large audio-language models (LALMs) exhibit strong zero-shot capabilities in multiple downstream tasks, such as audio question answering (AQA) and abstract reasoning; however, these models still lag behind specialized mo…

Audio ClassificationQuestion Answering