paper-with-me

홈 › Papers

Multi-modal Automated Speech Scoring using Attention Fusion

2020-05-17 · Manraj Singh Grover, Yaman Kumar, Sumit Sarin, Payman Vafaee, Mika Hama, Rajiv Ratn Shah

In this study, we propose a novel multi-modal end-to-end neural approach for automated assessment of non-native English speakers' spontaneous speech using attention fusion. The pipeline employs Bi-directional Recurrent Convolutional Neural Networks and Bi-directional Long Short-Term Memory Neural Networks to encode acoustic and lexical cues from spectrograms and transcriptions, respectively. Attention fusion is performed on these learned predictive features to learn complex interactions between different modalities before final scoring. We compare our model with strong baselines and find combined attention to both lexical and acoustic cues significantly improves the overall performance of the system. Further, we present a qualitative and quantitative analysis of our model.

📄 PDF Abstract BibTeX arXiv:2005.08182

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Pronunciation Assessment with Multi-modal Large Language Models

2024-07-12 · Kaiqi Fu, Linkai Peng, Nan Yang, Shuran Zhou

Large language models (LLMs), renowned for their powerful conversational abilities, are widely recognized as exceptional tools in the field of education, particularly in the context of automated intelligent instruction s…

My Teacher Thinks The World Is Flat! Interpreting Automatic Essay Scoring Mechanism

2020-12-27 · Swapnil Parekh, Yaman Kumar Singla, Changyou Chen, Junyi Jessy Li 외

Significant progress has been made in deep-learning based Automatic Essay Scoring (AES) systems in the past two decades. However, little research has been put to understand and interpret the black-box nature of these dee…

Common Sense ReasoningNatural Language UnderstandingWorld Knowledge

Speech Recognition Rescoring with Large Speech-Text Foundation Models

2024-09-25 · Prashanth Gurunath Shivakumar, Jari Kolehmainen, Aditya Gourav, Yi Gu 외

Large language models (LLM) have demonstrated the ability to understand human language by leveraging large amount of text data. Automatic speech recognition (ASR) systems are often limited by available transcribed speech…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+2

Using Ontology-based Approaches to Representing Speech Transcripts for Automated Speech Scoring

2012-06-01 · NAACL 2012 6 · Miao Chen
Speech Recognition

A Novel Data Augmentation Approach for Automatic Speaking Assessment on Opinion Expressions

2025-06-04 · Chung-Chun Wang, Jhen-Ke Lin, Hao-Chien Lu, Hong-Yun Lin 외

Automated speaking assessment (ASA) on opinion expressions is often hampered by the scarcity of labeled recordings, which restricts prompt diversity and undermines scoring reliability. To address this challenge, we propo…

Data AugmentationDiversityLanguage ModelingLanguage Modelling+6