paper-with-me

Papers

English Pronunciation Evaluation without Complex Joint Training: LoRA Fine-tuned Speech Multimodal LLM

2025-09-03 · Taekyung Ahn, Hosung Nam arxiv

This study demonstrates that a Multimodal Large Language Model (MLLM) adapted via Low-Rank Adaptation (LoRA) can perform both Automatic Pronunciation Assessment (APA) and Mispronunciation Detection and Diagnosis (MDD) simultaneously. Leveraging Microsoft's Phi-4-multimodal-instruct, our fine-tuning method eliminates the need for complex architectural changes or separate training procedures conventionally required for these distinct tasks. Fine-tuned on the Speechocean762 dataset, the pronunciation evaluation scores predicted by the model exhibited a strong Pearson Correlation Coefficient (PCC > 0.7) with human-assigned scores, while achieving low Word Error Rate (WER) and Phoneme Error Rate (PER) (both < 0.15). Notably, fine-tuning only the LoRA layers was sufficient to achieve performance levels comparable to those achieved by fine-tuning all audio layers. This research highlights that an integrated pronunciation assessment system can be established by adapting large multimodal models without full fine-tuning, utilizing a significantly simpler training methodology compared to previous joint models designed for simultaneous APA and MDD. This efficient LoRA-based approach paves the way for more accessible, integrated, and effective Computer-Assisted Pronunciation Training (CAPT) technologies for English L2 learners.

📄 PDF Abstract BibTeX arXiv:2509.02915

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

An Investigation of Indian Native Language Phonemic Influences on L2 English Pronunciations

2022-12-19 · Shelly Jain, Priyanshi Pal, Anil Vuppala, Prasanta Ghosh 외

Speech systems are sensitive to accent variations. This is especially challenging in the Indian context, with an abundance of languages but a dearth of linguistic studies characterising pronunciation variations. The grow…

Experiments of ASR-based mispronunciation detection for children and adult English learners

2021-04-13 · Nina Hosseini-Kivanani, Roberto Gretter, Marco Matassoni, Giuseppe Daniele Falavigna

Pronunciation is one of the fundamentals of language learning, and it is considered a primary factor of spoken language when it comes to an understanding and being understood by others. The persistent presence of high er…

Language Modellingspeech-recognitionSpeech Recognition

Assessment of an Index for Measuring Pronunciation Difficulty

2018-07-01 · WS 2018 7 · Katsunori Kotani, Takehiko Yoshimi

This study assesses an index for measur-ing the pronunciation difficulty of sen-tences (henceforth, pronounceability) based on the normalized edit distance from a reference sentence to a transcrip-tion of learners{'} pro…

Sentencevalid

AUTOMATIC PRONUNCIATION MISTAKE DETECTOR PROJECT REPORT

2025-06-25 · ResearchGate 2025 6 · Kamal Acharya

Given the drawbacks of traditional English pronunciation correction systems, such as failure to provide timely feedback and correct learners' pronunciation errors, slow improvement of learners' English proficiency, and e…

Mistake Detectionspeech-recognitionSpeech Recognition

Non-native English lexicon creation for bilingual speech synthesis

2021-06-21 · Arun Baby, Pranav Jawale, Saranya Vinnaitherthan, Sumukh Badam 외

Bilingual English speakers speak English as one of their languages. Their English is of a non-native kind, and their conversations are of a code-mixed fashion. The intelligibility of a bilingual text-to-speech (TTS) syst…

Speech Synthesistext-to-speechText to Speech