paper-with-me

Papers

Dolphin: A Large-Scale Automatic Speech Recognition Model for Eastern Languages

2025-03-26 · Yangyang Meng, Jinpeng Li, Guodong Lin, Yu Pu, Guanbo Wang, Hu Du, Zhiming Shao, YuKai Huang, Ke Li, Wei-Qiang Zhang

This report introduces Dolphin, a large-scale multilingual automatic speech recognition (ASR) model that extends the Whisper architecture to support a wider range of languages. Our approach integrates in-house proprietary and open-source datasets to refine and optimize Dolphin's performance. The model is specifically designed to achieve notable recognition accuracy for 40 Eastern languages across East Asia, South Asia, Southeast Asia, and the Middle East, while also supporting 22 Chinese dialects. Experimental evaluations show that Dolphin significantly outperforms current state-of-the-art open-source models across various languages. To promote reproducibility and community-driven innovation, we are making our trained models and inference source code publicly available.

📄 PDF Abstract BibTeX arXiv:2503.20212

Code (1)

dataoceanai/dolphin 공식 구현 pytorch

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local Attention

2025-09-28 · Kai Li, Kejun Gao, Xiaolin Hu arxiv

Audio-visual speech separation (AVSS) methods leverage visual cues to extract target speech and have demonstrated strong separation quality in noisy acoustic environments. However, these methods usually involve a large n…

Speech Separation

Dolphin: Closed-loop Open-ended Auto-research through Thinking, Practice, and Feedback

2025-01-07 · Jiakang Yuan, Xiangchao Yan, Botian Shi, Tao Chen 외

The scientific research paradigm is undergoing a profound transformation owing to the development of Artificial Intelligence (AI). Recent works demonstrate that various AI-assisted research methods can largely improve re…

image-classificationImage Classification

Dolphin-CN-Dialect: Where Chinese Dialects Matter

2026-05-09 · Yangyang Meng, Huihang Zhong, Guodong Lin, Guanbo Wang 외 arxiv

We present Dolphin-CN-Dialect, a streaming-capable ASR model with a focus on Chinese and dialect-rich scenarios. Compared to the previous version, Dolphin-CN-Dialect introduces substantial improvements in data processing…

Dolphin v1.0 Technical Report

2025-09-30 · Taohan Weng, Kaibing Hu, Henan Liu, Siya Liu 외 arxiv

Ultrasound is crucial in modern medicine but faces challenges like operator dependence, image noise, and real-time scanning, hindering AI integration. While large multimodal models excel in other medical imaging areas, t…

Reinforcement Learning

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting

2025-05-20 · Hao Feng, Shu Wei, Xiang Fei, Wei Shi 외

Document image parsing is challenging due to its complexly intertwined elements such as text paragraphs, figures, formulas, and tables. Current approaches either assemble specialized expert models or directly generate pa…