paper-with-me

홈 › Papers

SURE-Challenge: Evaluating Speech Evidence Before Speech-LLM Generation

2026-08-27 · Mengzhe Geng arxiv

Speech LLMs are usually graded after they answer, although an operating system first has to decide whether a waveform should be sent to the model. We define the Speech-Unsupported Rejection Evaluation Challenge (SURE-Challenge) for this admission step. The benchmark pairs LibriSpeech-derived transcription and first-word question answering with unsupported silence, colored noise, synthetic tones, and source-ambiguous babble under disjoint source splits. Front-end ablations use Qwen2-Audio; the selected energy-plus-Whisper-score rule is then replayed before six speech/audio LLMs. On the 474-row leakage-screened SURE-Extended test set, raw Qwen2-Audio rejects 15/204 unsupported inputs, whereas the fixed rule rejects 196/204 and leaves supported accuracy unchanged. External checks delimit this number: Common Voice retention drops as the Whisper-score threshold is tightened, and no-speed babble gives 18 to 24 rejected clips out of 54 across regenerated seeds. The result identifies a pre-generation error mode missed by answer-only scoring.

📄 PDF Abstract BibTeX arXiv:2608.27783

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Similar Papers 제목 키워드 기반

Evaluating Gender Bias in Speech Translation

2020-10-27 · LREC 2022 6 · Marta R. Costa-jussà, Christine Basta, Gerard I. Gállego

The scientific community is increasingly aware of the necessity to embrace pluralism and consistently represent major and minor social groups. Currently, there are no standard evaluation techniques for different types of…

Translation

Evaluating Speech Synthesis by Training Recognizers on Synthetic Speech

2023-10-01 · Dareen Alharthi, Roshan Sharma, Hira Dhamyal, Soumi Maiti 외

Modern speech synthesis systems have improved significantly, with synthetic speech being indistinguishable from real speech. However, efficient and holistic evaluation of synthetic speech still remains a significant chal…

speech-recognitionSpeech RecognitionSpeech Synthesistext-to-speech+1

Feature Selection Enhancement and Feature Space Visualization for Speech-Based Emotion Recognition

2022-08-19 · Sofia Kanwal, Sohail Asghar, Hazrat Ali

Robust speech emotion recognition relies on the quality of the speech features. We present speech features enhancement strategy that improves speech emotion recognition. We used the INTERSPEECH 2010 challenge feature-set…

Emotion Recognitionfeature selectionSpeech Emotion Recognition

Multi-dimensional Speech Quality Assessment in Crowdsourcing

2023-09-14 · Babak Naderi, Ross Cutler, Nicolae-Catalin Ristea

Subjective speech quality assessment is the gold standard for evaluating speech enhancement processing and telecommunication systems. The commonly used standard ITU-T Rec. P.800 defines how to measure speech quality in l…

Speech Enhancement

Speech Disfluencies occur at Higher Perplexities

2020-12-01 · COLING (CogALex) 2020 12 · Priyanka Sen

Speech disfluencies have been hypothesized to occur before words that are less predictable and therefore more cognitively demanding. In this paper, we revisit this hypothesis by using OpenAI’s GPT-2 to calculate predicta…

Language ModelingLanguage Modelling