paper-with-me

홈 › Papers

AudioBench: A Universal Benchmark for Audio Large Language Models

2024-06-23 · Bin Wang, Xunlong Zou, Geyu Lin, Shuo Sun, Zhuohan Liu, Wenyu Zhang, Zhengyuan Liu, AiTi Aw, Nancy F. Chen

We introduce AudioBench, a universal benchmark designed to evaluate Audio Large Language Models (AudioLLMs). It encompasses 8 distinct tasks and 26 datasets, among which, 7 are newly proposed datasets. The evaluation targets three main aspects: speech understanding, audio scene understanding, and voice understanding (paralinguistic). Despite recent advancements, there lacks a comprehensive benchmark for AudioLLMs on instruction following capabilities conditioned on audio signals. AudioBench addresses this gap by setting up datasets as well as desired evaluation metrics. Besides, we also evaluated the capabilities of five popular models and found that no single model excels consistently across all tasks. We outline the research outlook for AudioLLMs and anticipate that our open-sourced evaluation toolkit, data, and leaderboard will offer a robust testbed for future model developments.

📄 PDF Abstract BibTeX arXiv:2406.16020

Code (1)

audiollms/audiobench 공식 구현 pytorch

Tasks

Audio Scene UnderstandingInstruction FollowingScene Understanding

Similar Papers 제목 키워드 기반

ParaPairAudioBench: Paralinguistic Pairwise Audio Benchmark for LALM-as-a-Judge

2026-06-23 · Jisu Jeon, Seungyeon Jwa, Joosung Lee, Jinhyeon Kim 외 arxiv

Large Audio-Language Models (LALMs) have been widely used as judge models for the automatic evaluation of generated speech. However, prior approaches predominantly focus on holistic naturalness, leaving fine-grained para…

Jailbreak-AudioBench: In-Depth Evaluation and Analysis of Jailbreak Threats for Large Audio Language Models

2025-01-23 · Hao Cheng, Erjia Xiao, Jing Shao, Yichi Wang 외

Large Language Models (LLMs) demonstrate impressive zero-shot performance across a wide range of natural language processing tasks. Integrating various modality encoders further expands their capabilities, giving rise to…

Safety Alignment

KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs

2026-05-27 · Haechan Kim, Seungjun Chung, Inkyu Park, Jihoo Lee 외 arxiv

Speech language models (SpeechLMs) have achieved substantial progress by extending large language models (LLMs) to the speech modality. However, SpeechLM evaluation remains heavily centered on English, limiting reliable …

Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

2025-03-06 · Sreyan Ghosh, Zhifeng Kong, Sonal Kumar, S Sakshi 외

Understanding and reasoning over non-speech sounds and music are crucial for both humans and AI agents to interact effectively with their environments. In this paper, we introduce Audio Flamingo 2 (AF2), an Audio-Languag…

Audio captioningLanguage ModelingLanguage ModellingQuestion Answering+1

AC/DC: LLM-based Audio Comprehension via Dialogue Continuation

2025-06-12 · Yusuke Fujita, Tomoya Mizumoto, Atsushi Kojima, Lianbo Liu 외

We propose an instruction-following audio comprehension model that leverages the dialogue continuation ability of large language models (LLMs). Instead of directly generating target captions in training data, the propose…

AudioCapsAudio captioningInstruction FollowingQuestion Answering