paper-with-me

Papers

MERaLiON-SpeechEncoder: Towards a Speech Foundation Model for Singapore and Beyond

2024-12-16 · Muhammad Huzaifah, Geyu Lin, Tianchi Liu, Hardik B. Sailor, Kye Min Tan, Tarun K. Vangani, Qiongqiong Wang, Jeremy H. M. Wong, Nancy F. Chen, Ai Ti Aw

This technical report describes the MERaLiON-SpeechEncoder, a foundation model designed to support a wide range of downstream speech applications. Developed as part of Singapore's National Multimodal Large Language Model Programme, the MERaLiON-SpeechEncoder is tailored to address the speech processing needs in Singapore and the surrounding Southeast Asian region. The model currently supports mainly English, including the variety spoken in Singapore. We are actively expanding our datasets to gradually cover other languages in subsequent releases. The MERaLiON-SpeechEncoder was pre-trained from scratch on 200,000 hours of unlabelled speech data using a self-supervised learning approach based on masked language modelling. We describe our training procedure and hyperparameter tuning experiments in detail below. Our evaluation demonstrates improvements to spontaneous and Singapore speech benchmarks for speech recognition, while remaining competitive to other state-of-the-art speech encoders across ten other speech tasks. We commit to releasing our model, supporting broader research endeavours, both in Singapore and beyond.

📄 PDF Abstract BibTeX arXiv:2412.11538

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language ModelSelf-Supervised Learningspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

MERaLiON-AudioLLM: Bridging Audio and Language with Large Language Models

2024-12-13 · Yingxu He, Zhuohan Liu, Shuo Sun, Bin Wang 외

We introduce MERaLiON-AudioLLM (Multimodal Empathetic Reasoning and Learning in One Network), the first speech-text model tailored for Singapore's multilingual and multicultural landscape. Developed under the National La…

speech-recognitionSpeech Recognition

MERaLiON-SER: Robust Speech Emotion Recognition Model for English and SEA Languages

2025-11-07 · Hardik B. Sailor, Aw Ai Ti, Chen Fang Yih Nancy, Chiu Ying Lay 외 arxiv

We present MERaLiON-SER, a robust speech emotion recognition model designed for English and Southeast Asian languages. The model is trained using a hybrid objective combining weighted categorical cross-entropy and Concor…

Speech Emotion RecognitionMultimodal Reasoning

Polyglot-Lion: Efficient Multilingual ASR for Singapore via Balanced Fine-Tuning of Qwen3-ASR

2026-03-17 · Quy-Anh Dang, Chris Ngo arxiv

We present Polyglot-Lion, a family of compact multilingual automatic speech recognition (ASR) models tailored for the linguistic landscape of Singapore, covering English, Mandarin, Tamil, and Malay. Our models are obtain…

Speech Recognition

MF-Speech: Achieving Fine-Grained and Compositional Control in Speech Generation via Factor Disentanglement

2025-11-15 · Xinyue Yu, Youqing Fang, Pingyu Wu, Guoyang Ye 외 arxiv

Generating expressive and controllable human speech is one of the core goals of generative artificial intelligence, but its progress has long been constrained by two fundamental challenges: the deep entanglement of speec…

ASTRA: A Scalable Next-Generation ATCO Training Simulator with Autonomous Simpilots

2026-06-16 · Ethan Chew, Enjia Wu, Iruss Eng, Ian Lim 외 arxiv

Air Traffic Control Operators (ATCOs) are vital in ensuring the safe, orderly, and efficient flow of air traffic, yet training capacity is constrained by reliance on specialized human trainers known as simpilots, who mus…

Speech Recognition