paper-with-me

Papers

MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark

2025-06-05 · Dingdong Wang, Jincenzi Wu, Junan Li, Dongchao Yang, Xueyuan Chen, Tianhua Zhang, Helen Meng

Speech inherently contains rich acoustic information that extends far beyond the textual language. In real-world spoken language understanding, effective interpretation often requires integrating semantic meaning (e.g., content), paralinguistic features (e.g., emotions, speed, pitch) and phonological characteristics (e.g., prosody, intonation, rhythm), which are embedded in speech. While recent multimodal Speech Large Language Models (SpeechLLMs) have demonstrated remarkable capabilities in processing audio information, their ability to perform fine-grained perception and complex reasoning in natural speech remains largely unexplored. To address this gap, we introduce MMSU, a comprehensive benchmark designed specifically for understanding and reasoning in spoken language. MMSU comprises 5,000 meticulously curated audio-question-answer triplets across 47 distinct tasks. To ground our benchmark in linguistic theory, we systematically incorporate a wide range of linguistic phenomena, including phonetics, prosody, rhetoric, syntactics, semantics, and paralinguistics. Through a rigorous evaluation of 14 advanced SpeechLLMs, we identify substantial room for improvement in existing models, highlighting meaningful directions for future optimization. MMSU establishes a new standard for comprehensive assessment of spoken language understanding, providing valuable insights for developing more sophisticated human-AI speech interaction systems. MMSU benchmark is available at https://huggingface.co/datasets/ddwang2000/MMSU. Evaluation Code is available at https://github.com/dingdongwang/MMSU_Bench.

📄 PDF Abstract BibTeX arXiv:2506.04779

Code (2)

dingdongwang/mmsu_bench 공식 구현
QwenLM/Qwen2.5-Omni pytorch

Tasks

RhythmSpoken Language Understanding

Similar Papers 제목 키워드 기반

MMSummary: Multimodal Summary Generation for Fetal Ultrasound Video

2024-08-07 · Xiaoqing Guo, Qianhui Men, J. Alison Noble

We present the first automated multimodal summary generation system, MMSummary, for medical imaging video, particularly with a focus on fetal ultrasound analysis. Imitating the examination process performed by a human so…

AnatomyLanguage ModelingLanguage ModellingLarge Language Model

StepAudio 3 Realtime Technical Report

2026-09-12 · Bin Lin, Bo Zhao, Boyang Zhang, Boyong Wu 외 hf

Realtime spoken interaction demands deep reasoning, prompt responses, and fluid turn-taking. We present StepAudio 3 Realtime, an audio-language foundation model organized around a continuous listen-converse-think-act loo…

MMSum: A Dataset for Multimodal Summarization and Thumbnail Generation of Videos

2023-06-07 · CVPR 2024 1 · JieLin Qiu, Jiacheng Zhu, William Han, Aditesh Kumar 외

Multimodal summarization with multimodal output (MSMO) has emerged as a promising research direction. Nonetheless, numerous limitations exist within existing public MSMO datasets, including insufficient maintenance, data…

Text SummarizationVideo Summarization

Speech-MASSIVE: A Multilingual Speech Dataset for SLU and Beyond

2024-08-07 · Beomseok Lee, Ioan Calapodescu, Marco Gaido, Matteo Negri 외

We present Speech-MASSIVE, a multilingual Spoken Language Understanding (SLU) dataset comprising the speech counterpart for a portion of the MASSIVE textual corpus. Speech-MASSIVE covers 12 languages from different famil…

BenchmarkingLanguage Identificationslot-fillingSlot Filling+1

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment

2026-07-23 · Dongjie Fu, Di Cao, Xize Cheng, Zihan Zhang 외 arxiv

While large audio-language models have achieved remarkable progress in auditory perception, they still lag behind text-based large language models in deep logical reasoning, primarily due to the scarcity of high-quality …

Logical Reasoning