paper-with-me

홈 › Papers

Uncertainty Modeling in Multimodal Speech Analysis Across the Psychosis Spectrum

2025-02-25 · Morteza Rohanian, Roya M. Hüppi, Farhad Nooralahzadeh, Noemi Dannecker, Yves Pauli, Werner Surbeck, Iris Sommer, Wolfram Hinzen, Nicolas Langer, Michael Krauthammer, Philipp Homan

Capturing subtle speech disruptions across the psychosis spectrum is challenging because of the inherent variability in speech patterns. This variability reflects individual differences and the fluctuating nature of symptoms in both clinical and non-clinical populations. Accounting for uncertainty in speech data is essential for predicting symptom severity and improving diagnostic precision. Speech disruptions characteristic of psychosis appear across the spectrum, including in non-clinical individuals. We develop an uncertainty-aware model integrating acoustic and linguistic features to predict symptom severity and psychosis-related traits. Quantifying uncertainty in specific modalities allows the model to address speech variability, improving prediction accuracy. We analyzed speech data from 114 participants, including 32 individuals with early psychosis and 82 with low or high schizotypy, collected through structured interviews, semi-structured autobiographical tasks, and narrative-driven interactions in German. The model improved prediction accuracy, reducing RMSE and achieving an F1-score of 83% with ECE = 4.5e-2, showing robust performance across different interaction contexts. Uncertainty estimation improved model interpretability by identifying reliability differences in speech markers such as pitch variability, fluency disruptions, and spectral instability. The model dynamically adjusted to task structures, weighting acoustic features more in structured settings and linguistic features in unstructured contexts. This approach strengthens early detection, personalized assessment, and clinical decision-making in psychosis-spectrum research.

📄 PDF Abstract BibTeX arXiv:2502.18285

Code (0)

등록된 구현이 없습니다.

Tasks

Diagnostic

Similar Papers 제목 키워드 기반

Modeling speech emotion with label variance and analyzing performance across speakers and unseen acoustic conditions

2025-03-24 · Vikramjit Mitra, Amrit Romana, Dung T. Tran, Erdrin Azemi

Spontaneous speech emotion data usually contain perceptual grades where graders assign emotion score after listening to the speech files. Such perceptual grades introduce uncertainty in labels due to grader opinion varia…

Emotion RecognitionOverall - Test

Label Uncertainty Modeling and Prediction for Speech Emotion Recognition using t-Distributions

2022-07-25 · Navin Raj Prabhu, Nale Lehmann-Willenbrock, Timo Gerkmann

As different people perceive others' emotional expressions differently, their annotation in terms of arousal and valence are per se subjective. To address this, these emotion annotations are typically collected by multip…

Emotion RecognitionSpeech Emotion Recognition

Data-Centric Improvements for Enhancing Multi-Modal Understanding in Spoken Conversation Modeling

2024-12-20 · Maximillian Chen, Ruoxi Sun, Sercan Ö. Arik

Conversational assistants are increasingly popular across diverse real-world applications, highlighting the need for advanced multimodal speech modeling. Speech, as a natural mode of communication, encodes rich user-spec…

Multi-Task Learning

Uncertainty-Aware Multimodal Emotion Recognition through Dirichlet Parameterization

2026-02-09 · Rémi Grzeczkowicz, Eric Soriano, Ali Janati, Miyu Zhang 외 arxiv

In this work, we present a lightweight and privacy-preserving Multimodal Emotion Recognition (MER) framework designed for deployment on edge devices. To demonstrate framework's versatility, our implementation uses three …

Multimodal Emotion Recognition

Integrating Uncertainty into Neural Network-based Speech Enhancement

2023-05-15 · Huajian Fang, Dennis Becker, Stefan Wermter, Timo Gerkmann

Supervised masking approaches in the time-frequency domain aim to employ deep neural networks to estimate a multiplicative mask to extract clean speech. This leads to a single estimate for each input without any guarante…

Speech Enhancement