paper-with-me

Papers

Optimizing Speech-Input Length for Speaker-Independent Depression Classification

2024-12-31 · Tomasz Rutowski, Amir Harati, Yang Lu, Elizabeth Shriberg

Machine learning models for speech-based depression classification offer promise for health care applications. Despite growing work on depression classification, little is understood about how the length of speech-input impacts model performance. We analyze results for speaker-independent depression classification using a corpus of over 1400 hours of speech from a human-machine health screening application. We examine performance as a function of response input length for two NLP systems that differ in overall performance. Results for both systems show that performance depends on natural length, elapsed length, and ordering of the response within a session. Systems share a minimum length threshold, but differ in a response saturation threshold, with the latter higher for the better system. At saturation it is better to pose a new question to the speaker, than to continue the current response. These and additional reported results suggest how applications can be better designed to both elicit and process optimal input lengths for depression classification.

📄 PDF Abstract BibTeX arXiv:2501.00608

Code (0)

등록된 구현이 없습니다.

Tasks

Classification

Similar Papers 제목 키워드 기반

Exploring Timbre Disentanglement in Non-Autoregressive Cross-Lingual Text-to-Speech

2021-10-14 · Haoyue Zhan, Xinyuan Yu, Haitong Zhang, Yang Zhang 외

In this paper, we study the disentanglement of speaker and language representations in non-autoregressive cross-lingual TTS models from various aspects. We propose a phoneme length regulator that solves the length mismat…

Disentanglementtext-to-speechText to SpeechVoice Cloning

Discriminative Learning for Monaural Speech Separation Using Deep Embedding Features

2019-07-23 · Cunhang Fan, Bin Liu, Jian-Hua Tao, Jiangyan Yi 외

Deep clustering (DC) and utterance-level permutation invariant training (uPIT) have been demonstrated promising for speaker-independent speech separation. DC is usually formulated as two-step processes: embedding learnin…

ClusteringDeep ClusteringSpeech Separation

A Novel Speech Feature Fusion Algorithm for Text-Independent Speaker Recognition

2022-12-01 · Biao Ma, Chengben Xu, Ye Zhang

A novel speech feature fusion algorithm with independent vector analysis (IVA) and parallel convolutional neural network (PCNN) is proposed for text-independent speaker recognition. Firstly, some different feature types,…

Speaker RecognitionText-Independent Speaker Recognition

Removing Speaker Information from Speech Representation using Variable-Length Soft Pooling

2024-04-01 · Injune Hwang, Kyogu Lee

Recently, there have been efforts to encode the linguistic information of speech using a self-supervised framework for speech synthesis. However, predicting representations from surrounding representations can inadverten…

Speaker IdentificationSpeech Synthesis

Spatial Pyramid Encoding with Convex Length Normalization for Text-Independent Speaker Verification

2019-06-19 · Youngmoon Jung, Younggwan Kim, Hyungjun Lim, Yeunju Choi 외

In this paper, we propose a new pooling method called spatial pyramid encoding (SPE) to generate speaker embeddings for text-independent speaker verification. We first partition the output feature maps from a deep residu…

Speaker VerificationText-Independent Speaker Verification