paper-with-me

홈 › Papers

ISA-Bench: Benchmarking Instruction Sensitivity for Large Audio Language Models

2025-10-27 · Bohan Li, Wenbin Huang, Yuhang Qiu, Yiwei Guo, Hankun Wang, Zhihan Li, Jing Peng, Ziyang Ma, Xie Chen, Kai Yu arxiv

Large Audio Language Models (LALMs), which couple acoustic perception with large language models (LLMs) to extract and understand diverse information from audio, have attracted intense interest from both academic and industrial communities. However, existing LALMs are highly sensitive to how instructions are phrased, affecting both (i) instruction-following rates and (ii) task performance. Yet, no existing benchmarks offer a systematic and comprehensive evaluation of this sensitivity. We introduce ISA-Bench, a dynamic benchmark evaluating instruction sensitivity for LALMs along three axes: instruction description, output format, and task composition. We assess recent open-source and proprietary LALMs using ISA-Bench, profiling both compliance and accuracy under controlled instruction variations. Experimental results reveal that even state-of-the-art LALMs suffer significant instruction sensitivity, leading to degraded performance on fundamental audio understanding tasks. To mitigate this issue, we fine-tune Qwen2-Audio on a specifically constructed complex instruction-variant dataset, achieving a marked improvement in instruction-following performance. However, this also induces nontrivial catastrophic forgetting: the model loses some previously mastered task capabilities when exposed to new instruction styles. Our benchmark provides a standardized basis for assessing and improving instruction sensitivity in LALMs, underscoring the need for instruction-robust audio understanding in real-world pipelines.

📄 PDF Abstract BibTeX arXiv:2510.23558

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

IFEval-Audio: Benchmarking Instruction-Following Capability in Audio-based Large Language Models

2025-05-22 · Yiming Gao, Bin Wang, Chengwei Wei, Shuo Sun 외

Large language models (LLMs) have demonstrated strong instruction-following capabilities in text-based tasks. However, this ability often deteriorates in multimodal models after alignment with non-text modalities such as…

BenchmarkingInstruction Following

OmniCap-IF: Benchmarking and Improving Instruction Following Abilities for Omni-Video Captioning

2026-06-07 · Jiahao Wang, An Ping, Yanghai Wang, Yuanxing Zhang 외 arxiv

While Omni-modal Large Language Models (OLLMs) have demonstrated impressive capabilities in jointly processing audio and visual streams, their ability to strictly adhere to complex, multi-faceted user instructions remain…

Instruction FollowingVideo Captioning

JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models

2025-05-23 · Zifan Peng, Yule Liu, Zhen Sun, Mingchen Li 외

Audio Language Models (ALMs) have made significant progress recently. These models integrate the audio modality directly into the model, rather than converting speech into text and inputting text to Large Language Models…

BenchmarkingDiversity

AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension

2024-02-12 · Qian Yang, Jin Xu, Wenrui Liu, Yunfei Chu 외

Recently, instruction-following audio-language models have received broad attention for human-audio interaction. However, the absence of benchmarks capable of evaluating audio-centric interaction capabilities has impeded…

2kAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Benchmarking+3

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models

2026-07-15 · Chun-Yi Kuan, Siwon Kim, Byeonggeun Kim, Suyoun Kim 외 arxiv

Recent text-to-audio models generate high-quality audio, but often fail to follow instructions involving multiple sound events and temporal order. This gap arises because existing evaluation and training signals mainly e…

Instruction Following