paper-with-me

Papers

When LLMs Meets Acoustic Landmarks: An Efficient Approach to Integrate Speech into Large Language Models for Depression Detection

2024-02-17 · Xiangyu Zhang, Hexin Liu, Kaishuai Xu, Qiquan Zhang, Daijiao Liu, Beena Ahmed, Julien Epps

Depression is a critical concern in global mental health, prompting extensive research into AI-based detection methods. Among various AI technologies, Large Language Models (LLMs) stand out for their versatility in mental healthcare applications. However, their primary limitation arises from their exclusive dependence on textual input, which constrains their overall capabilities. Furthermore, the utilization of LLMs in identifying and analyzing depressive states is still relatively untapped. In this paper, we present an innovative approach to integrating acoustic speech information into the LLMs framework for multimodal depression detection. We investigate an efficient method for depression detection by integrating speech signals into LLMs utilizing Acoustic Landmarks. By incorporating acoustic landmarks, which are specific to the pronunciation of spoken words, our method adds critical dimensions to text transcripts. This integration also provides insights into the unique speech patterns of individuals, revealing the potential mental states of individuals. Evaluations of the proposed approach on the DAIC-WOZ dataset reveal state-of-the-art results when compared with existing Audio-Text baselines. In addition, this approach is not only valuable for the detection of depression but also represents a new perspective in enhancing the ability of LLMs to comprehend and process speech signals.

📄 PDF Abstract BibTeX arXiv:2402.13276

Code (0)

등록된 구현이 없습니다.

Tasks

Depression Detection

Similar Papers 제목 키워드 기반

When CTC Training Meets Acoustic Landmarks

2018-11-05 · Di He, Xuesong Yang, Boon Pang Lim, Yi Liang 외

Connectionist temporal classification (CTC) provides an end-to-end acoustic model (AM) training strategy. CTC learns accurate AMs without time-aligned phonetic transcription, but sometimes fails to converge, especially i…

Automatic Speech Recognition (ASR)Speech Recognition

Auto-Landmark: Acoustic Landmark Dataset and Open-Source Toolkit for Landmark Extraction

2024-09-12 · Xiangyu Zhang, Daijiao Liu, Tianyi Xiao, Cihan Xiao 외

In the speech signal, acoustic landmarks identify times when the acoustic manifestations of the linguistically motivated distinctive features are most salient. Acoustic landmarks have been widely applied in various domai…

Depression Detectionspeech-recognitionSpeech Recognition

Generating Talking Face Landmarks from Speech

2018-03-26 · Sefik Emre Eskimez, Ross K. Maddox, Chenliang Xu, Zhiyao Duan

The presence of a corresponding talking face has been shown to significantly improve speech intelligibility in noisy conditions and for hearing impaired population. In this paper, we present a system that can generate la…

Improved ASR for Under-Resourced Languages Through Multi-Task Learning with Acoustic Landmarks

2018-05-15 · Di He, Boon Pang Lim, Xuesong Yang, Mark Hasegawa-Johnson 외

Furui first demonstrated that the identity of both consonant and vowel can be perceived from the C-V transition; later, Stevens proposed that acoustic landmarks are the primary cues for speech perception, and that steady…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Multi-Task Learningspeech-recognition+1

Detection of Acoustic-Phonetic Landmarks in Mismatched Conditions using a Biomimetic Model of Human Auditory Processing

2012-12-01 · COLING 2012 12 · Sarah King, Mark Hasegawa-Johnson
Speech Recognition