paper-with-me

Papers

K-Function: Joint Pronunciation Transcription and Feedback for Evaluating Kids Language Function

2025-07-03 · Shuhe Li, Chenxu Guo, Jiachen Lian, Cheol Jun Cho, Wenshuo Zhao, Xiner Xu, Ruiyu Jin, Xiaoyu Shi, Xuanru Zhou, Dingkun Zhou, Sam Wang, Grace Wang, Jingze Yang, Jingyi Xu, Ruohan Bao, Xingrui Chen, Elise Brenner, Brandon In, Francesca Pei, Maria Luisa Gorno-Tempini, Gopala Anumanchipalli arxiv

Evaluating young children's language is challenging for automatic speech recognizers due to high-pitched voices, prolonged sounds, and limited data. We introduce K-Function, a framework that combines accurate sub-word transcription with objective, Large Language Model (LLM)-driven scoring. Its core, Kids-Weighted Finite State Transducer (K-WFST), merges an acoustic phoneme encoder with a phoneme-similarity model to capture child-specific speech errors while remaining fully interpretable. K-WFST achieves a 1.39 % phoneme error rate on MyST and 8.61 % on Multitudes-an absolute improvement of 10.47 % and 7.06 % over a greedy-search decoder. These high-quality transcripts are used by an LLM to grade verbal skills, developmental milestones, reading, and comprehension, with results that align closely with human evaluators. Our findings show that precise phoneme recognition is essential for creating an effective assessment framework, enabling scalable language screening for children.

📄 PDF Abstract BibTeX arXiv:2507.03043

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

QVoice: Arabic Speech Pronunciation Learning Application

2023-05-09 · Yassine El Kheir, Fouad Khnaisser, Shammur Absar Chowdhury, Hamdy Mubarak 외

This paper introduces a novel Arabic pronunciation learning application QVoice, powered with end-to-end mispronunciation detection and feedback generator module. The application is designed to support non-native Arabic s…

Basis Identification for Automatic Creation of Pronunciation Lexicon for Proper Names

2014-06-05 · Sunil Kumar Kopparapu, M Laxminarayana

Development of a proper names pronunciation lexicon is usually a manual effort which can not be avoided. Grapheme to phoneme (G2P) conversion modules, in literature, are usually rule based and work best for non-proper na…

Automatic Speech Recognition (ASR) for the Diagnosis of pronunciation of Speech Sound Disorders in Korean children

2024-03-13 · Taekyung Ahn, Yeonjung Hong, Younggon Im, Do Hyung Kim 외

This study presents a model of automatic speech recognition (ASR) designed to diagnose pronunciation issues in children with speech sound disorders (SSDs) to replace manual transcriptions in clinical procedures. Since AS…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Diagnosticspeech-recognition+1

Automatic Pronunciation Generation by Utilizing a Semi-supervised Deep Neural Networks

2016-06-15 · Naoya Takahashi, Tofigh Naghibi, Beat Pfister

Phonemic or phonetic sub-word units are the most commonly used atomic elements to represent speech signals in modern ASRs. However they are not the optimal choice due to several reasons such as: large amount of effort re…

speech-recognitionSpeech Recognition

Exploring the Potential of Large Multimodal Models as Effective Alternatives for Pronunciation Assessment

2025-03-14 · Ke Wang, Lei He, Kun Liu, Yan Deng 외

Large Multimodal Models (LMMs) have demonstrated exceptional performance across a wide range of domains. This paper explores their potential in pronunciation assessment tasks, with a particular focus on evaluating the ca…