paper-with-me

Papers

M3-SLU: Evaluating Speaker-Attributed Reasoning in Multimodal Large Language Models

2025-10-22 · Yejin Kwon, Taewoo Kang, Hyunsoo Yoon, Changouk Kim arxiv

We present M3-SLU, a new multimodal large language model (MLLM) benchmark for evaluating multi-speaker, multi-turn spoken language understanding. While recent models show strong performance in speech and text comprehension, they still struggle with speaker-attributed reasoning, the ability to understand who said what and when in natural conversations. M3-SLU is built from four open corpora (CHiME-6, MELD, MultiDialog, and AMI) and comprises over 12,000 validated instances with paired audio, transcripts, and metadata. It includes two tasks: (1) Speaker-Attributed Question Answering and (2) Speaker Attribution via Utterance Matching. We provide baseline results for both cascaded pipelines and end-to-end MLLMs, evaluated using an LLM-as-Judge and accuracy metrics. Results show that while models can capture what was said, they often fail to identify who said it, revealing a key gap in speaker-aware dialogue understanding. M3-SLU offers as a challenging benchmark to advance research in speaker-aware multimodal understanding.

📄 PDF Abstract BibTeX arXiv:2510.19358

Code (0)

등록된 구현이 없습니다.

Tasks

Spoken Language UnderstandingQuestion Answering

Similar Papers 제목 키워드 기반

MOSS Transcribe Diarize Technical Report

2026-01-04 · MOSI. AI, :, Donghua Yu, Zhengyuan Lin 외 arxiv

Speaker-Attributed, Time-Stamped Transcription (SATS) aims to transcribe what is said and to precisely determine the timing of each speaker, which is particularly valuable for meeting transcription. Existing SATS systems…

See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models

2025-12-01 · Le Thien Phuc Nguyen, Zhuoran Yu, Samuel Low Yu Hang, Subin An 외 arxiv

Multimodal large language models (MLLMs) are expected to jointly interpret vision, audio, and language, yet existing video benchmarks rarely assess fine-grained reasoning about human speech. Many tasks remain visually so…

Speaker-Reasoner: Scaling Interaction Turns and Reasoning Patterns for Timestamped Speaker-Attributed ASR

2026-04-03 · Zhennan Lin, Shuai Wang, Zhaokai Sun, Pengyuan Xie 외 arxiv

Transcribing and understanding multi-speaker conversations requires speech recognition, speaker attribution, and timestamp localization. While speech LLMs excel at single-speaker tasks, multi-speaker scenarios remain cha…

Speech Recognition

Real-TurnTurk: A Multimodal Turkish Corpus for Turn-Taking Prediction

2026-08-22 · Ahmet Tuğrul Bayrak, Fatma Nur Korkmaz, Bekir Berker Türker, Mustafa Sertaç Türkel 외 hf

Turn-taking is a basic organizational feature of human conversation and remains difficult to model in natural, synchronous dialog systems. While existing research has explored multimodal approaches and large language mod…

Binary Classification

Minimum Bayes Risk Training for End-to-End Speaker-Attributed ASR

2020-11-03 · Naoyuki Kanda, Zhong Meng, Liang Lu, Yashesh Gaur 외

Recently, an end-to-end speaker-attributed automatic speech recognition (E2E SA-ASR) model was proposed as a joint model of speaker counting, speech recognition and speaker identification for monaural overlapped speech. …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speaker Identificationspeech-recognition+1