paper-with-me

Papers

Fine-Tuning Large Multimodal Models for Automatic Pronunciation Assessment

2025-09-19 · Ke Wang, Wenning Wei, Yan Deng, Lei He, Sheng Zhao arxiv

Automatic Pronunciation Assessment (APA) is critical for Computer-Assisted Language Learning (CALL), requiring evaluation across multiple granularities and aspects. Large Multimodal Models (LMMs) present new opportunities for APA, but their effectiveness in fine-grained assessment remains uncertain. This work investigates fine-tuning LMMs for APA using the Speechocean762 dataset and a private corpus. Fine-tuning significantly outperforms zero-shot settings and achieves competitive results on single-granularity tasks compared to public and commercial systems. The model performs well at word and sentence levels, while phoneme-level assessment remains challenging. We also observe that the Pearson Correlation Coefficient (PCC) reaches 0.9, whereas Spearman's rank Correlation Coefficient (SCC) remains around 0.6, suggesting that SCC better reflects ordinal consistency. These findings highlight both the promise and limitations of LMMs for APA and point to future work on fine-grained modeling and rank-aware evaluation.

📄 PDF Abstract BibTeX arXiv:2509.15701

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

English Pronunciation Evaluation without Complex Joint Training: LoRA Fine-tuned Speech Multimodal LLM

2025-09-03 · Taekyung Ahn, Hosung Nam arxiv

This study demonstrates that a Multimodal Large Language Model (MLLM) adapted via Low-Rank Adaptation (LoRA) can perform both Automatic Pronunciation Assessment (APA) and Mispronunciation Detection and Diagnosis (MDD) si…

Fine-Tuning Self-Supervised Learning Models for End-to-End Pronunciation Scoring

2023-09-19 · IEEE Access 2023 9 · Ahmed I. Zahran, Aly A. Fahmy, Khaled T. Wassif, Hanaa Bayomi

Automatic pronunciation assessment models are regularly used in language learning applications. Common methodologies for pronunciation assessment use feature-based approaches, such as the Goodness-of-Pronunciation (GOP) …

Feature EngineeringPhone-level pronunciation scoringPhoneme RecognitionSelf-Supervised Learning+3

Real-time Ultrasound-enhanced Multimodal Imaging of Tongue using 3D Printable Stabilizer System: A Deep Learning Approach

2019-11-22 · M. Hamed Mozaffari, Won-Sook Lee

Despite renewed awareness of the importance of articulation, it remains a challenge for instructors to handle the pronunciation needs of language learners. There are relatively scarce pedagogical tools for pronunciation …

Causal Structure Discovery for Error Diagnostics of Children's ASR

2025-05-31 · Vishwanath Pratap Singh, Md. Sahidullah, Tomi Kinnunen

Children's automatic speech recognition (ASR) often underperforms compared to that of adults due to a confluence of interdependent factors: physiological (e.g., smaller vocal tracts), cognitive (e.g., underdeveloped pron…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Neural TTS in French: Comparing Graphemic and Phonetic Inputs Using the SynPaFlex-Corpus and Tacotron2

2023-04-17 · Samuel Delalez, Ludi Akue

The SynPaFlex-Corpus is a publicly available TTS-oriented dataset, which provides phonetic transcriptions automatically produced by the JTrans transcriber, with a Phoneme Error Rate (PER) of 6.1%. In this paper, we analy…