paper-with-me

Papers

Detection of Consonant Errors in Disordered Speech Based on Consonant-vowel Segment Embedding

2021-06-16 · Si-Ioi Ng, Cymie Wing-Yee Ng, Jingyu Li, Tan Lee

Speech sound disorder (SSD) refers to a type of developmental disorder in young children who encounter persistent difficulties in producing certain speech sounds at the expected age. Consonant errors are the major indicator of SSD in clinical assessment. Previous studies on automatic assessment of SSD revealed that detection of speech errors concerning short and transitory consonants is less satisfactory. This paper investigates a neural network based approach to detecting consonant errors in disordered speech using consonant-vowel (CV) diphone segment in comparison to using consonant monophone segment. The underlying assumption is that the vowel part of a CV segment carries important information of co-articulation from the consonant. Speech embeddings are extracted from CV segments by a recurrent neural network model. The similarity scores between the embeddings of the test segment and the reference segments are computed to determine if the test segment is the expected consonant or not. Experimental results show that using CV segments achieves improved performance on detecting speech errors concerning those "difficult" consonants reported in the previous studies.

📄 PDF Abstract BibTeX arXiv:2106.08536

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

1x1 Convolution A 1 x 1 Convolution is a convolution with some special properties in that it can be used for dimensionality reduction,…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Non Maximum Suppression Non Maximum Suppression is a computer vision method that selects a single entity out of many overlapping entities (for example bounding boxes in object detection). The…
SSD SSD is a single-stage object detection method that discretizes the output space of bounding boxes into a set of default boxes over different aspect ratios and scales per…

Similar Papers 제목 키워드 기반

Finding My Voice: Generative Reconstruction of Disordered Speech for Automated Clinical Evaluation

2025-09-23 · Karen Rosero, Eunjung Yeo, David R. Mortensen, Cortney Van't Slot 외 arxiv

We present ChiReSSD, a speech reconstruction framework that preserves children speaker's identity while suppressing mispronunciations. Unlike prior approaches trained on healthy adult speech, ChiReSSD adapts to the voice…

Automatic Estimation of Intelligibility Measure for Consonants in Speech

2020-05-12 · Ali Abavisani, Mark Hasegawa-Johnson

In this article, we provide a model to estimate a real-valued measure of the intelligibility of individual speech segments. We trained regression models based on Convolutional Neural Networks (CNN) for stop consonants \t…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Persian Vowel recognition with MFCC and ANN on PCVC speech dataset

2018-12-17 · Saber Malekzadeh, Mohammad Hossein Gholizadeh, Seyed Naser Razavi

In this paper a new method for recognition of consonant-vowel phonemes combination on a new Persian speech dataset titled as PCVC (Persian Consonant-Vowel Combination) is proposed which is used to recognize Persian phone…

Phoneme Recognition

The role of vowel and consonant onsets in neural tracking of natural speech

2023-07-31 · Mohammad Jalilpour Monesi, Jonas Vanthornhout, Hugo Van hamme, Tom Francart

To investigate how the auditory system processes natural speech, models have been created to relate the electroencephalography (EEG) signal of a person listening to speech to various representations of the speech. Mainly…

EEG

Consonant-Vowel Transition Models Based on Deep Learning for Objective Evaluation of Articulation

2022-03-18 · Vikram C. Mathad, Julie M. Liss, Kathy Chapman, Nancy Scherer 외

Spectro-temporal dynamics of consonant-vowel (CV) transition regions are considered to provide robust cues related to articulation. In this work, we propose an objective measure of precise articulation, dubbed the object…