paper-with-me

Papers

SpeechR: A Benchmark for Speech Reasoning in Large Audio-Language Models

2025-08-04 · Wanqi Yang, Yanda Li, Yunchao Wei, Meng Fang, Ling Chen arxiv

Large audio-language models (LALMs) have achieved near-human performance in sentence-level transcription and emotion recognition. However, existing evaluations focus mainly on surface-level perception, leaving the capacity of models for contextual and inference-driven reasoning in speech-based scenarios insufficiently examined. To address this gap, we introduce SpeechR, a unified benchmark for evaluating reasoning over speech in large audio-language models. SpeechR evaluates models along three key dimensions: factual retrieval, procedural inference, and normative judgment. It includes three distinct evaluation formats. The multiple-choice version measures answer selection accuracy. The generative version assesses the coherence and logical consistency of reasoning chains. The acoustic-feature version investigates whether variations in stress and emotion affect reasoning performance. Evaluations on eleven state-of-the-art LALMs reveal that high transcription accuracy does not translate into strong reasoning capabilities. SpeechR establishes a structured benchmark for evaluating reasoning in spoken language, enabling more targeted analysis of model capabilities across diverse dialogue-based tasks.

📄 PDF Abstract BibTeX arXiv:2508.02018

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion RecognitionAnswer Selection

Similar Papers 제목 키워드 기반

SpeechRole: A Large-Scale Dataset and Benchmark for Evaluating Speech Role-Playing Agents

2025-08-04 · Changhao Jiang, Jiajun Sun, Yifei Cao, Jiabao Zhuang 외 arxiv

Speech is essential for realistic role-playing, yet existing work on role-playing agents largely centers on text, leaving Speech Role-Playing Agents (SRPAs) underexplored and without systematic evaluation. We introduce S…

SpeechRefiner: Towards Perceptual Quality Refinement for Front-End Algorithms

2025-06-16 · Sirui Li, Shuai Wang, Zhijun Liu, Zhongjie Jiang 외

Speech pre-processing techniques such as denoising, de-reverberation, and separation, are commonly employed as front-ends for various downstream speech processing tasks. However, these methods can sometimes be inadequate…

Denoising

Development and Evaluation of Video Recordings for the OLSA Matrix Sentence Test

2019-12-10 · Gerard Llorach, Frederike Kirschner, Giso Grimm, Melanie A. Zokoll 외

One of the established multi-lingual methods for testing speech intelligibility is the matrix sentence test (MST). Most versions of this test are designed with audio-only stimuli. Nevertheless, visual cues play an import…

Sentence

CommonVoice-SpeechRE and RPG-MoGe: Advancing Speech Relation Extraction with a New Dataset and Multi-Order Generative Framework

2025-09-10 · Jinzhong Ning, Paerhati Tulajiang, Yingying Le, Yijia Zhang 외 arxiv

Speech Relation Extraction (SpeechRE) aims to extract relation triplets directly from speech. However, existing benchmark datasets rely heavily on synthetic data, lacking sufficient quantity and diversity of real human s…

Relation Extraction

I Speak and You Find: Robust 3D Visual Grounding with Noisy and Ambiguous Speech Inputs

2025-06-17 · Yu Qi, Lipeng Gu, Honghua Chen, Liangliang Nan 외

Existing 3D visual grounding methods rely on precise text prompts to locate objects within 3D scenes. Speech, as a natural and intuitive modality, offers a promising alternative. Real-world speech inputs, however, often …

3D visual groundingContrastive LearningSpeech-to-TextVisual Grounding