paper-with-me

Papers

Massive Sound Embedding Benchmark (MSEB)

2026-02-06 · Georg Heigold, Ehsan Variani, Tom Bagby, Cyril Allauzen, Ji Ma, Shankar Kumar, Michael Riley arxiv

Audio is a critical component of multimodal perception, and any truly intelligent system must demonstrate a wide range of auditory capabilities. These capabilities include transcription, classification, retrieval, reasoning, segmentation, clustering, reranking, and reconstruction. Fundamentally, each task involves transforming a raw audio signal into a meaningful 'embedding' - be it a single vector, a sequence of continuous or discrete representations, or another structured form - which then serves as the basis for generating the task's final response. To accelerate progress towards robust machine auditory intelligence, we present the Massive Sound Embedding Benchmark (MSEB): an extensible framework designed to evaluate the auditory components of any multimodal system. In its first release, MSEB offers a comprehensive suite of eight core tasks, with more planned for the future, supported by diverse datasets, including the new, large-scale Simple Voice Questions (SVQ) dataset. Our initial experiments establish clear performance headrooms, highlighting the significant opportunity to improve real-world multimodal experiences where audio is a core signal. We encourage the research community to use MSEB to assess their algorithms and contribute to its growth. The library is publicly hosted at github.

📄 PDF Abstract BibTeX arXiv:2602.07143

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Benchmarking LLMs on the Massive Sound Embedding Benchmark (MSEB)

2026-05-06 · Cyril Allauzen, Tom Bagby, Georg Heigold, Ehsan Variani 외 arxiv

The Massive Sound Embedding Benchmark (MSEB) has emerged as a standard for evaluating the functional breadth of audio models. While initial baselines focused on specialized encoders, the shift toward "audio-native" Large…

MAEB: Massive Audio Embedding Benchmark

2026-02-17 · Adnan El Assadi, Isaac Chung, Chenghao Xiao, Roman Solomatin 외 arxiv

We introduce the Massive Audio Embedding Benchmark (MAEB), a large-scale benchmark covering 30 tasks across speech, music, environmental sounds, and cross-modal audio-text reasoning in 100+ languages. We evaluate 50+ mod…

Environmental Sound Classification

SMRC: Aligning Large Language Models with Student Reasoning for Mathematical Error Correction

2025-11-18 · Biaojie Zeng, Min Zhang, Juan Zhou, Fengrui Liu 외 arxiv

Large language models (LLMs) often make reasoning errors when solving mathematical problems, and how to automatically detect and correct these errors has become an important research direction. However, existing approach…

High School Mathematics

SoundNet: Learning Sound Representations from Unlabeled Video

2016-10-27 · NeurIPS 2016 12 · Yusuf Aytar, Carl Vondrick, Antonio Torralba

We learn rich natural sound representations by capitalizing on large amounts of unlabeled sound data collected in the wild. We leverage the natural synchronization between vision and sound to learn an acoustic representa…

General Classification

The Computation of Generalized Embeddings for Underwater Acoustic Target Recognition using Contrastive Learning

2025-05-19 · Hilde I. Hummel, Arwin Gansekoele, Sandjai Bhulai, Rob van der Mei

The increasing level of sound pollution in marine environments poses an increased threat to ocean health, making it crucial to monitor underwater noise. By monitoring this noise, the sources responsible for this pollutio…

Contrastive LearningSound Classification