paper-with-me

홈 › Papers

Optimizing Automatic Speech Assessment: W-RankSim Regularization and Hybrid Feature Fusion Strategies

2024-06-16 · Chung-Wen Wu, Berlin Chen

Automatic Speech Assessment (ASA) has seen notable advancements with the utilization of self-supervised features (SSL) in recent research. However, a key challenge in ASA lies in the imbalanced distribution of data, particularly evident in English test datasets. To address this challenge, we approach ASA as an ordinal classification task, introducing Weighted Vectors Ranking Similarity (W-RankSim) as a novel regularization technique. W-RankSim encourages closer proximity of weighted vectors in the output layer for similar classes, implying that feature vectors with similar labels would be gradually nudged closer to each other as they converge towards corresponding weighted vectors. Extensive experimental evaluations confirm the effectiveness of our approach in improving ordinal classification performance for ASA. Furthermore, we propose a hybrid model that combines SSL and handcrafted features, showcasing how the inclusion of handcrafted features enhances performance in an ASA system.

📄 PDF Abstract BibTeX arXiv:2406.10873

Code (0)

등록된 구현이 없습니다.

Tasks

Ordinal Classification

Similar Papers 제목 키워드 기반

RankSim: Ranking Similarity Regularization for Deep Imbalanced Regression

2022-05-30 · Yu Gong, Greg Mori, Frederick Tung

Data imbalance, in which a plurality of the data samples come from a small proportion of labels, poses a challenge in training deep neural networks. Unlike classification, in regression the labels are continuous, potenti…

Deep imbalanced regressionInductive BiasregressionSTS+1

Automatic Severity Classification of Dysarthric speech by using Self-supervised Model with Multi-task Learning

2022-10-27 · Eun Jung Yeo, Kwanghee Choi, Sunhee Kim, Minhwa Chung

Automatic assessment of dysarthric speech is essential for sustained treatments and rehabilitation. However, obtaining atypical speech is challenging, often leading to data scarcity issues. To tackle the problem, we prop…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)ClassificationMulti-Task Learning+2

Suicide Risk Assessment Using Multimodal Speech Features: A Study on the SW1 Challenge Dataset

2025-05-19 · Ambre Marie, Ilias Maoudj, Guillaume Dardenne, Gwenolé Quellec

The 1st SpeechWellness Challenge conveys the need for speech-based suicide risk assessment in adolescents. This study investigates a multimodal approach for this challenge, integrating automatic transcription with Whispe…

Comparison of End-to-end Speech Assessment Models for the NOCASA 2025 Challenge

2025-09-03 · Aleksei Žavoronkov, Tanel Alumäe arxiv

This paper presents an analysis of three end-to-end models developed for the NOCASA 2025 Challenge, aimed at automatic word-level pronunciation assessment for children learning Norwegian as a second language. Our models …

Speech Quality Assessment Model Based on Mixture of Experts: System-Level Performance Enhancement and Utterance-Level Challenge Analysis

2025-07-08 · Xintong Hu, Yixuan Chen, Rui Yang, Wenxiang Guo 외

Automatic speech quality assessment plays a crucial role in the development of speech synthesis systems, but existing models exhibit significant performance variations across different granularity levels of prediction ta…

Data AugmentationMixture-of-ExpertsPredictionSelf-Supervised Learning+5