paper-with-me

Papers

FRICATIVE PHONEME DETECTION WITH ZERO DELAY

2019-09-25 · Metehan Yurt, Alberto N. Escalante B., Veniamin I. Morgenshtern

People with high-frequency hearing loss rely on hearing aids that employ frequency lowering algorithms. These algorithms shift some of the sounds from the high frequency band to the lower frequency band where the sounds become more perceptible for the people with the condition. Fricative phonemes have an important part of their content concentrated in high frequency bands. It is important that the frequency lowering algorithm is activated exactly for the duration of a fricative phoneme, and kept off at all other times. Therefore, timely (with zero delay) and accurate fricative phoneme detection is a key problem for high quality hearing aids. In this paper we present a deep learning based fricative phoneme detection algorithm that has zero detection delay and achieves state-of-the-art fricative phoneme detection accuracy on the TIMIT Speech Corpus. All reported results are reproducible and come with easy to use code that could serve as a baseline for future research.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Phoneme-Level Deepfake Detection Across Emotional Conditions Using Self-Supervised Embeddings

2026-05-04 · Vamshi Nallaguntla, Shruti Kshirsagar, Anderson R. Avila arxiv

Recent advances in emotional voice conversion (EVC) have enabled the generation of expressive synthetic speech, raising new concerns in audio deepfake detection. Existing approaches treat speech as a homogeneous signal a…

Audio Deepfake DetectionVoice Conversion

Improving Word Recognition in Speech Transcriptions by Decision-level Fusion of Stemming and Two-way Phoneme Pruning

2021-07-26 · Sunakshi Mehra, Seba Susan

We introduce an unsupervised approach for correcting highly imperfect speech transcriptions based on a decision-level fusion of stemming and two-way phoneme pruning. Transcripts are acquired from videos by extracting aud…

VoXtream: Full-Stream Text-to-Speech with Extremely Low Latency

2025-09-19 · Nikita Torgashov, Gustav Eje Henter, Gabriel Skantze arxiv

We present VoXtream, a fully autoregressive, zero-shot streaming text-to-speech (TTS) system for real-time use that begins speaking from the first word. VoXtream directly maps incoming phonemes to audio tokens using a mo…

Benchmarking Multilingual Speech Models on Pashto: Zero-Shot ASR, Script Failure, and Cross-Domain Evaluation

2026-04-06 · Hanif Rahman arxiv

Pashto is spoken by approximately 60--80 million people but has no published benchmarks for multilingual automatic speech recognition (ASR) on any shared public test set. This paper reports the first reproducible multi-m…

Speech Recognition

MARTA: a model for the automatic phonemic grouping of the parkinsonian speech

2024-03-19 · techrxiv 2024 3 · Alejandro Guerrero-López, Julián D. Arias-Londoño, Stefanie Shattuck-Hufnagel, Juan I. Godino-Llorente

Parkinson's disease significantly impacts speech, particularly affecting phonemic groups like stop-plosives, fricatives, and affricates. However, its objective impact on the different phonemic groups has been briefly add…

BenchmarkingClassificationMetric Learning