paper-with-me

홈 › Papers

Do SpeechLMs Hear Their Own Opinions? Diagnosing and Mitigating Previous-Belief Contamination in Streaming Emotion Understanding

2026-08-21 · Haoyue Liu, Zhichao Wang, Ye Chen, Haonan Deng, Xiaoying Tang arxiv

Streaming emotion understanding uses historical state while continuously interpreting current audio, often feeding the model's previous prediction back as context. We show that this history conditioning can distort current perception. On a balanced CREMA-D-Stream counterfactual diagnostic, changing only the injected previous emotion label while holding the audio fixed reduces current-audio accuracy from 72.50% to 30.42% and flips 65.69% of predictions. The effect is strongly label-asymmetric, with prior pull ranging from 4.76% to 98.20%, revealing a failure we call previous-belief contamination (PBC). To address PBC, we introduce EmoUpdate, a training-free framework that separates current-audio perception from historical state revision through three components: (1) a prior-blind acoustic firewall that prevents historical state from entering perception; (2) an evidence-shrunk causal belief filter that introduces history only after observation formation and retains label-asymmetric transition structure only when supported by observed evidence; and (3) a closed-form decontamination operator derived from the same counterfactual measurements for serving stacks where firewalling is unavailable. Across four SpeechLMs and two streaming emotion benchmarks, EmoUpdate achieves the best step accuracy and state-balanced accuracy in all eight model--benchmark settings, improving S-BAcc by up to 69.71 points and step accuracy by up to 38.41 points over the strongest controlled baselines.

📄 PDF Abstract BibTeX arXiv:2608.20769

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

HearHere: Mitigating Echo Chambers in News Consumption through an AI-based Web System

2024-02-28 · Youngseung Jeon, JaeHoon Kim, SoHyun Park, Yunyong Ko 외

Considerable efforts are currently underway to mitigate the negative impacts of echo chambers, such as increased susceptibility to fake news and resistance towards accepting scientific evidence. Prior research has presen…

LoASR-Bench: Evaluating Large Speech Language Models on Low-Resource Automatic Speech Recognition Across Language Families

2026-03-20 · Jianan Chen, Xiaoxue Gao, Tatsuya Kawahara, Nancy F. Chen arxiv

Large language models (LLMs) have driven substantial advances in speech language models (SpeechLMs), yielding strong performance in automatic speech recognition (ASR) under high-resource conditions. However, existing ben…

Speech Recognition

Enhancing Speech Large Language Models through Reinforced Behavior Alignment

2025-08-25 · Yansong Liu, Jiateng Li, Yuan Liu arxiv

The recent advancements of Large Language Models (LLMs) have spurred considerable research interest in extending their linguistic capabilities beyond text to other modalities, which leads to emergence of speech-based LLM…

Speech-to-Text TranslationReinforcement LearningQuestion Answering

Monitoring of the heart movements using a FMCW radar and correlation with an ECG

2023-01-21 · Rémi Grisot, Pierre Laurent, Claire Migliaccio, Jean-Yves Dauvignac 외

Monitoring the activity of the heart is important for diagnosing and preventing cardiovascular diseases. The electrocardiogram (ECG) is the gold standard for diagnosing such diseases. It monitors the heart's electrical a…

Recent Advances in Speech Language Models: A Survey

2024-10-01 · Wenqian Cui, Dianzhi Yu, Xiaoqi Jiao, Ziqiao Meng 외

Large Language Models (LLMs) have recently garnered significant attention, primarily for their capabilities in text-based interactions. However, natural human interaction often relies on speech, necessitating a shift tow…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+3