paper-with-me

홈 › Papers

CASA: Content-Acoustic Speaking Assessment with Speech Encoder and Large Language Model

2026-08-13 · Nhan Phan, Ilona Lähteenmäki, Anna von Zansen, Olli-Pekka Pauna, Yaroslav Getman, Tamás Grósz, Mikko Kurimo arxiv

Research on automatic speaking assessment (ASA) has increasingly adopted multimodal speech large language models to assess learners' speaking performance. However, existing studies provide limited analysis of how acoustic and content information contribute to predictions and how stable the resulting performance is. We propose CASA, a simpler architecture combining Whisper-medium and Qwen3.5-2B that achieves state-of-the-art performance while providing a more interpretable separation between speech delivery and content. On the Speak & Improve Corpus 2025, CASA achieves a root mean square error (RMSE) of 0.358, improving on the previous best RMSE while using approximately half the estimated inference parameters. The general-purpose architecture is designed for adaptation to other ASA corpora without structural changes and relies on three handcrafted fluency features. Through ablations and repeated runs, we analyze the individual and complementary contributions of acoustic and content information, examine performance variability, and demonstrate the potential of large language model reasoning for training-free content validation.

📄 PDF Abstract BibTeX arXiv:2608.13101

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

An Effective Strategy for Modeling Score Ordinality and Non-uniform Intervals in Automated Speaking Assessment

2025-08-27 · Tien-Hong Lo, Szu-Yu Chen, Yao-Ting Sung, Berlin Chen arxiv

A recent line of research on automated speaking assessment (ASA) has benefited from self-supervised learning (SSL) representations, which capture rich acoustic and linguistic patterns in non-native speech without underly…

Self-Supervised Learning

Addressing Cold Start Problem for End-to-end Automatic Speech Scoring

2023-06-25 · Jungbae Park, Seungtaek Choi

Integrating automatic speech scoring/assessment systems has become a critical aspect of second-language speaking education. With self-supervised learning advancements, end-to-end speech scoring approaches have exhibited …

Self-Supervised Learning

Non-native Children's Automatic Speech Assessment Challenge (NOCASA)

2025-04-29 · Yaroslav Getman, Tamás Grósz, Mikko Kurimo, Giampiero Salvi

This paper presents the "Non-native Children's Automatic Speech Assessment" (NOCASA) - a data competition part of the IEEE MLSP 2025 conference. NOCASA challenges participants to develop new systems that can assess singl…

Investigating Human-Model Discrepancies in Speech Quality Assessment via Acoustic and Prosodic Perturbations

2026-06-18 · Masato Takagi, Masaya Kawamura, Reo Shimizu, Yuma Shirahata arxiv

Mean opinion score (MOS) prediction models are widely used as proxy metrics in text-to-speech (TTS) research, yet their ability to capture quality differences beyond acoustic fidelity remains unclear. We investigate this…

Beyond Modality Limitations: A Unified MLLM Approach to Automated Speaking Assessment with Effective Curriculum Learning

2025-08-18 · Yu-Hsuan Fang, Tien-Hong Lo, Yao-Ting Sung, Berlin Chen arxiv

Traditional Automated Speaking Assessment (ASA) systems exhibit inherent modality limitations: text-based approaches lack acoustic information while audio-based methods miss semantic context. Multimodal Large Language Mo…