paper-with-me

Papers

MAEB: Massive Audio Embedding Benchmark

2026-02-17 · Adnan El Assadi, Isaac Chung, Chenghao Xiao, Roman Solomatin, Animesh Jha, Rahul Chand, Silky Singh, Kaitlyn Wang, Ali Sartaz Khan, Marc Moussa Nasser, Sufen Fong, Pengfei He, Alan Xiao, Ayush Sunil Munot, Aditya Shrivastava, Artem Gazizov, Niklas Muennighoff, Kenneth Enevoldsen arxiv

We introduce the Massive Audio Embedding Benchmark (MAEB), a large-scale benchmark covering 30 tasks across speech, music, environmental sounds, and cross-modal audio-text reasoning in 100+ languages. We evaluate 50+ models and find that no single model dominates across all tasks: contrastive audio-text models excel at environmental sound classification (e.g., ESC50) but score near random on multilingual speech tasks (e.g., SIB-FLEURS), while speech-pretrained models show the opposite pattern. Clustering remains challenging for all models, with even the best-performing model achieving only modest results. We observe that models excelling on acoustic understanding often perform poorly on linguistic tasks, and vice versa. We also show that the performance of audio encoders on MAEB correlates highly with their performance when used in audio large language models. MAEB is derived from MAEB+, a collection of 98 tasks. MAEB is designed to maintain task diversity while reducing evaluation cost, and it integrates into the MTEB ecosystem for unified evaluation across text, image, and audio modalities. We release MAEB and all 98 tasks along with code and a leaderboard at https://github.com/embeddings-benchmark/mteb.

📄 PDF Abstract BibTeX arXiv:2602.16008

Code (0)

등록된 구현이 없습니다.

Tasks

Environmental Sound Classification

Similar Papers 제목 키워드 기반

Conan-embedding-v3: Fusing Modality-Specific Models for Omni-Modal Embedding

2026-06-08 · Shiyu Li, Zhiyuan Hu, Yifan Wang, Peiming Li 외 arxiv

Omni-modal retrieval promises a single embedding space for text, image, video, document, and audio inputs, but building such a unified retriever is difficult since these modalities differ in data distribution, architectu…

Omni-Interactive Universal Embedder

2026-08-27 · Wei-Yao Wang, Kazuya Tateishi, Shuyang Cui, Christian Simon 외 arxiv

Multimodal representation learning has been shifting from traditional two-tower architectures to large language model (LLM)-based embedders due to their strong instruction-following capabilities. Despite this progress, e…

Representation Learning

Obstacle detection test in real-word traffic contexts for the purposes of motorcycle autonomous emergency braking (MAEB)

2017-06-25 · Giovanni Savino, Simone Piantini, Gustavo Gil, Marco Pierini

Research suggests that a Motorcycle Autonomous Emergency Braking system (MAEB) could influence 25% of the crashes involving powered two wheelers (PTWs). By automatically slowing down a host PTW of up to 10 km/h in inevit…

MAEBE: Multi-Agent Emergent Behavior Framework

2025-06-03 · Sinem Erisken, Timothy Gothard, Martin Leitgab, Ram Potham

Traditional AI safety evaluations on isolated LLMs are insufficient as multi-agent AI ensembles become prevalent, introducing novel emergent risks. This paper introduces the Multi-Agent Emergent Behavior Evaluation (MAEB…

MVEB: Massive Video Embedding Benchmark

2026-06-12 · Adnan El Assadi, Roman Solomatin, Isaac Chung, Chenghao Xiao 외 arxiv

We introduce the Massive Video Embedding Benchmark (MVEB), a 23-task benchmark for video embeddings spanning classification, zero-shot classification, clustering, pair classification, retrieval, and video-centric questio…

Question Answering