paper-with-me

홈 › Papers

Robustness assessment of large audio language models in multiple-choice evaluation

2025-10-06 · Fernando López, Santosh Kesiraju, Jordi Luque arxiv

Recent advances in large audio language models (LALMs) have primarily been assessed using a multiple-choice question answering (MCQA) framework. However, subtle changes, such as shifting the order of choices, result in substantially different results. Existing MCQA frameworks do not account for this variability and report a single accuracy number per benchmark or category. We dive into the MCQA evaluation framework and conduct a systematic study spanning three benchmarks (MMAU, MMAR and MMSU) and four models: Audio Flamingo 2, Audio Flamingo 3, Qwen2.5-Omni-7B-Instruct, and Kimi-Audio-7B-Instruct. Our findings indicate that models are sensitive not only to the ordering of choices, but also to the paraphrasing of the question and the choices. Finally, we propose a simpler evaluation protocol and metric that account for subtle variations and provide a more detailed evaluation report of LALMs within the MCQA framework.

📄 PDF Abstract BibTeX arXiv:2510.04584

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Similar Papers 제목 키워드 기반

Evaluating Semantic Fragility in Text-to-Audio Generation Systems Under Controlled Prompt Perturbations

2026-03-14 · Jiahui Wu arxiv

Recent advances in text-to-audio generation enable models to translate natural-language descriptions into diverse musical output. However, the robustness of these systems under semantically equivalent prompt variations r…

Semantic SimilarityAudio Generation

AudioTrust: Benchmarking the Multifaceted Trustworthiness of Audio Large Language Models

2025-05-22 · Kai Li, Can Shen, Yile Liu, Jirui Han 외

The rapid advancement and expanding applications of Audio Large Language Models (ALLMs) demand a rigorous understanding of their trustworthiness. However, systematic research on evaluating these models, particularly conc…

BenchmarkingFairnessHallucination

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment

2025-01-30 · Yuqin Cao, Xiongkuo Min, Yixuan Gao, Wei Sun 외

Many video-to-audio (VTA) methods have been proposed for dubbing silent AI-generated videos. An efficient quality assessment method for AI-generated audio-visual content (AGAV) is crucial for ensuring audio-visual qualit…

AudioJailbreak: Jailbreak Attacks against End-to-End Large Audio-Language Models

2025-05-20 · Guangke Chen, Fu Song, Zhe Zhao, Xiaojun Jia 외

Jailbreak attacks to Large audio-language models (LALMs) are studied recently, but they achieve suboptimal effectiveness, applicability, and practicability, particularly, assuming that the adversary can fully manipulate …

text-to-speechText to Speech

Enabling Auditory Large Language Models for Automatic Speech Quality Evaluation

2024-09-25 · Siyin Wang, Wenyi Yu, Yudong Yang, Changli Tang 외

Speech quality assessment typically requires evaluating audio from multiple aspects, such as mean opinion score (MOS) and speaker similarity (SIM) \etc., which can be challenging to cover using one small model designed f…

text-to-speechText to Speech