paper-with-me

홈 › Papers

HARNESS: Lightweight Distilled Arabic Speech Foundation Models

2026-03-31 · Vrunda N. Sukhadia, Shammur Absar Chowdhury arxiv

Large self-supervised speech (SSL) models achieve strong downstream performance, but their size limits deployment in resource-constrained settings. We present HArnESS, an Arabic-centric self-supervised speech model family trained from scratch with iterative self-distillation, together with lightweight student variants that offer strong accuracy-efficiency trade-offs on Automatic Speech Recognition (ASR), Dialect Identification (DID), and Speech Emotion Recognition (SER). Our approach begins with a large bilingual Arabic-English teacher and progressively distills its knowledge into compressed student models while preserving Arabic-relevant acoustic and paralinguistic representations. We further study PCA-based compression of the teacher supervision signal to better match the capacity of shallow and thin students. Compared with HuBERT and XLS-R, HArnESS consistently improves performance on Arabic downstream tasks, while the compressed models remain competitive under substantial structural reduction. These results position HArnESS as a practical and accessible Arabic-centric SSL foundation for real-world speech applications.

📄 PDF Abstract BibTeX arXiv:2604.14186

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Emotion RecognitionSpeech Recognition

Similar Papers 제목 키워드 기반

HARNESS: Lightweight Distilled Arabic Speech Foundation Models

2025-09-18 · Vrunda N. sukhadia, Shammur Absar Chowdhury arxiv

Large pre-trained speech models excel in downstream tasks but their deployment is impractical for resource-limited environments. In this paper, we introduce HArnESS, the first Arabic-centric self-supervised speech model …

Emotion Recognition

Toward a Web-based Speech Corpus for Algerian Dialectal Arabic Varieties

2017-04-01 · WS 2017 4 · Soumia Bougrine, Aicha Chorana, Abdallah Lakhdari, Hadda Cherroun

The success of machine learning for automatic speech processing has raised the need for large scale datasets. However, collecting such data is often a challenging task as it implies significant investment involving time …

Speech RecognitionSpeech Synthesis

ArabEmoNet: A Lightweight Hybrid 2D CNN-BiLSTM Model with Attention for Robust Arabic Speech Emotion Recognition

2025-09-01 · Ali Abouzeid, Bilal Elbouardi, Mohamed Maged, Shady Shehata arxiv

Speech emotion recognition is vital for human-computer interaction, particularly for low-resource languages like Arabic, which face challenges due to limited data and research. We introduce ArabEmoNet, a lightweight arch…

Speech Emotion Recognition

Distilled Feature Fields Enable Few-Shot Language-Guided Manipulation

2023-07-27 · William Shen, Ge Yang, Alan Yu, Jansen Wong 외

Self-supervised and language-supervised image models contain rich knowledge of the world that is important for generalization. Many robotic tasks, however, require a detailed understanding of 3D geometry, which is often …

3D geometryFew-Shot LearningLanguage ModelingLanguage Modelling

Towards stable AI systems for Evaluating Arabic Pronunciations

2025-08-27 · Hadi Zaatiti, Hatem Hajri, Osama Abdullah, Nader Masmoudi arxiv

Modern Arabic ASR systems such as wav2vec 2.0 excel at word- and sentence-level transcription, yet struggle to classify isolated letters. In this study, we show that this phoneme-level task, crucial for language learning…