paper-with-me

홈 › Papers

GeHirNet: A Gender-Aware Hierarchical Model for Voice Pathology Classification

2025-08-02 · Fan Wu, Kaicheng Zhao, Elgar Fleisch, Filipe Barata arxiv

AI-based voice analysis shows promise for disease diagnostics, but existing classifiers often fail to accurately identify specific pathologies because of gender-related acoustic variations and the scarcity of data for rare diseases. We propose a novel two-stage framework that first identifies gender-specific pathological patterns using ResNet-50 on Mel spectrograms, then performs gender-conditioned disease classification. We address class imbalance through multi-scale resampling and time warping augmentation. Evaluated on a merged dataset from four public repositories, our two-stage architecture with time warping achieves state-of-the-art performance (97.63\% accuracy, 95.25\% MCC), with a 5\% MCC improvement over single-stage baseline. This work advances voice pathology classification while reducing gender bias through hierarchical modeling of vocal characteristics.

📄 PDF Abstract BibTeX arXiv:2508.01172

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Voice Cloning for Dysarthric Speech Synthesis: Addressing Data Scarcity in Speech-Language Pathology

2025-03-03 · Birger Moell, Fredrik Sand Aronsson

This study explores voice cloning to generate synthetic speech replicating the unique patterns of individuals with dysarthria. Using the TORGO dataset, we address data scarcity and privacy challenges in speech-language p…

Speech SynthesisVoice Cloning

Neutral TTS Female Voice Corpus in Brazilian Portuguese

2023-10-08 · XLI Simpósio Brasileiro de Telecomunicações e Processamento de Sinais (SBrT2023) 2023 10 · Pedro H. L. Leite, Edmundo Hoyle, Álvaro Antelo, Luiz F. Kruszielski 외

This paper introduces a new dataset designed to address the limitations in high-quality, diverse and representative datasets for training text-to-speech (TTS) models, specifically for female voices in Brazilian Portugues…

Speech Synthesistext-to-speechText to SpeechTransfer Learning

Generating Multilingual Gender-Ambiguous Text-to-Speech Voices

2022-11-01 · Konstantinos Markopoulos, Georgia Maniati, Georgios Vamvoukakis, Nikolaos Ellinas 외

The gender of any voice user interface is a key element of its perceived identity. Recently, there has been increasing interest in interfaces where the gender is ambiguous rather than clearly identifying as female or mal…

text-to-speechText to Speech

Can AI Have a Personality? Prompt Engineering for AI Personality Simulation: A Chatbot Case Study in Gender-Affirming Voice Therapy Training

2025-08-20 · Tailon D. Jackson, Byunggu Yu arxiv

This thesis investigates whether large language models (LLMs) can be guided to simulate a consistent personality through prompt engineering. The study explores this concept within the context of a chatbot designed for Sp…

Prompt Engineering

Towards an Interpretable Representation of Speaker Identity via Perceptual Voice Qualities

2023-10-04 · Robin Netzorg, Bohan Yu, Andrea Guzman, Peter Wu 외

Unlike other data modalities such as text and vision, speech does not lend itself to easy interpretation. While lay people can understand how to describe an image or sentence via perception, non-expert descriptions of sp…

Sentence