paper-with-me

Papers

Pronunciation Assessment with Multi-modal Large Language Models

2024-07-12 · Kaiqi Fu, Linkai Peng, Nan Yang, Shuran Zhou

Large language models (LLMs), renowned for their powerful conversational abilities, are widely recognized as exceptional tools in the field of education, particularly in the context of automated intelligent instruction systems for language learning. In this paper, we propose a scoring system based on LLMs, motivated by their positive impact on text-related scoring tasks. Specifically, the speech encoder first maps the learner's speech into contextual features. The adapter layer then transforms these features to align with the text embedding in latent space. The assessment task-specific prefix and prompt text are embedded and concatenated with the features generated by the modality adapter layer, enabling the LLMs to predict accuracy and fluency scores. Our experiments demonstrate that the proposed scoring systems achieve competitive results compared to the baselines on the Speechocean762 datasets. Moreover, we also conducted an ablation study to better understand the contributions of the prompt text and training strategy in the proposed scoring system.

📄 PDF Abstract BibTeX arXiv:2407.09209

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Adapter 설명 없음
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Exploring the Potential of Large Multimodal Models as Effective Alternatives for Pronunciation Assessment

2025-03-14 · Ke Wang, Lei He, Kun Liu, Yan Deng 외

Large Multimodal Models (LMMs) have demonstrated exceptional performance across a wide range of domains. This paper explores their potential in pronunciation assessment tasks, with a particular focus on evaluating the ca…

English Pronunciation Evaluation without Complex Joint Training: LoRA Fine-tuned Speech Multimodal LLM

2025-09-03 · Taekyung Ahn, Hosung Nam arxiv

This study demonstrates that a Multimodal Large Language Model (MLLM) adapted via Low-Rank Adaptation (LoRA) can perform both Automatic Pronunciation Assessment (APA) and Mispronunciation Detection and Diagnosis (MDD) si…

Fine-Tuning Large Multimodal Models for Automatic Pronunciation Assessment

2025-09-19 · Ke Wang, Wenning Wei, Yan Deng, Lei He 외 arxiv

Automatic Pronunciation Assessment (APA) is critical for Computer-Assisted Language Learning (CALL), requiring evaluation across multiple granularities and aspects. Large Multimodal Models (LMMs) present new opportunitie…

MultiPA: A Multi-task Speech Pronunciation Assessment Model for Open Response Scenarios

2023-08-24 · Yu-Wen Chen, Zhou Yu, Julia Hirschberg

Pronunciation assessment models designed for open response scenarios enable users to practice language skills in a manner similar to real-life communication. However, previous open-response pronunciation assessment model…

Multi-Task LearningSentence

Transformer-Based Multi-Aspect Multi-Granularity Non-Native English Speaker Pronunciation Assessment

2022-05-06 · Yuan Gong, Ziyi Chen, Iek-Heng Chu, Peng Chang 외

Automatic pronunciation assessment is an important technology to help self-directed language learners. While pronunciation quality has multiple aspects including accuracy, fluency, completeness, and prosody, previous eff…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Multi-Task LearningPhone-level pronunciation scoring+4