paper-with-me

홈 › Papers

StressTest: Can YOUR Speech LM Handle the Stress?

2025-05-28 · Iddo Yosha, Gallil Maimon, Yossi Adi

Sentence stress refers to emphasis, placed on specific words within a spoken utterance to highlight or contrast an idea, or to introduce new information. It is often used to imply an underlying intention that is not explicitly stated. Recent advances in speech-aware language models (SLMs) have enabled direct processing of audio, allowing models to bypass transcription and access the full richness of the speech signal and perform audio reasoning tasks such as spoken question answering. Despite the crucial role of sentence stress in shaping meaning and speaker intent, it remains largely overlooked in evaluation and development of such models. In this work, we address this gap by introducing StressTest, a benchmark specifically designed to evaluate a model's ability to distinguish between interpretations of spoken sentences based on the stress pattern. We assess the performance of several leading SLMs and find that, despite their overall capabilities, they perform poorly on such tasks. To overcome this limitation, we propose a novel synthetic data generation pipeline, and create Stress17k, a training set that simulates change of meaning implied by stress variation. Then, we empirically show that optimizing models with this synthetic dataset aligns well with real-world recordings and enables effective finetuning of SLMs. Results suggest, that our finetuned model, StresSLM, significantly outperforms existing models on both sentence stress reasoning and detection tasks. Code, models, data, and audio samples - pages.cs.huji.ac.il/adiyoss-lab/stresstest.

📄 PDF Abstract BibTeX arXiv:2505.22765

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringSentenceSynthetic Data Generation

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Stress-Testing Multimodal Foundation Models for Crystallographic Reasoning

2025-06-16 · Can Polat, Hasan Kurban, Erchin Serpedin, Mustafa Kurban

Evaluating foundation models for crystallographic reasoning requires benchmarks that isolate generalization behavior while enforcing physical constraints. This work introduces a multiscale multicrystal dataset with two p…

HallucinationSpatial Interpolation

A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness

2026-02-04 · Leo Schwinn, Moritz Ladenburger, Tim Beyer, Mehrnaz Mofakhami 외 arxiv

Automated \enquote{LLM-as-a-Judge} frameworks have become the de facto standard for scalable evaluation across natural language processing. For instance, in safety evaluation, these judges are relied upon to evaluate har…

Adversarial Robustness

Towards the Development of Speech-Based Measures of Stress Response in Individuals

2021-06-01 · NAACL (CLPsych) 2021 6 · Archna Bhatia, Toshiya Miyatsu, Peter Pirolli

Psychological and physiological stress in the environment can induce a different stress response in different individuals. Given the causal relationship between stress, mental health, and psychopathologies, as well as it…

Automatic Detection of Stress from Speech in the Trier Social Stress Test

2026-07-01 · Hanna Drimalla, Wieland R. Cremer, Christine Kraus, Oliver T. Wolf arxiv

Automatically detecting stress in speech provides an unobtrusive way to gain insights relevant to behavioral research or clinical assessment. This study investigates the automatic differentiation between a stressful and …

Speaker Diarization

Adversarial speech for voice privacy protection from Personalized Speech generation

2024-01-22 · Shihao Chen, Liping Chen, Jie Zhang, KongAik Lee 외

The rapid progress in personalized speech generation technology, including personalized text-to-speech (TTS) and voice conversion (VC), poses a challenge in distinguishing between generated and real speech for human list…

Speaker Verificationtext-to-speechText to SpeechVoice Conversion