paper-with-me

Papers

Raon-Speech Technical Report

2026-04-08 · Beomsoo Kim, Changho Choi, Dohyun Kim, Dongki Lee, Ethan Ewer, Eunchong Kim, Gyeongman Kim, Haechan Kim, Hyeonghwan Kim, Inkyu Park, Jihun Yun, Jihwan Moon, Jiyun Kim, Joonghyun Bae, Junhyuck Kim, Minkyu Kim, Sehun Lee, Seungjun Chung, Sungwoo Cho, Dongmin Park, Dongwon Kim, Hara Kang, Jonghyun Lee, Keon Lee, Kangwook Lee, Jaewoong Cho arxiv

We present Raon-Speech, a top-performing 9B-parameter speech language model (SpeechLM) for English and Korean speech understanding, answering, and generation, and Raon-SpeechChat, a high-performing full-duplex extension for natural real-time conversation. Raon-Speech successfully transforms a pre-trained LLM into a SpeechLM that both understands and generates speech while preserving strong text capabilities. It trains on 1.38M hours of highly curated English and Korean speech and text datasets with the following training stages: (1) speech modules alignment, (2) end-to-end SpeechLM pre-training with knowledge distillation, and (3) multi-task preference optimization-based post-training. Across 42 English and Korean speech and text benchmarks, Raon-Speech establishes the strongest overall profile on speech-centric tasks in our comparison against eight similarly sized recent audio foundation models, including Qwen2.5-Omni and Fun-Audio-Chat, while preserving strong text question answering performance. Building upon it, Raon-SpeechChat enables natural full-duplex conversation by continual training on 119K hours of time-aligned real and synthetic dialogue data. It proceeds through three complementary training stages: (1) causal encoder adaptation, (2) full-duplex pre-training, (3) full-duplex fine-tuning for voice and role-control. On multiple full-duplex benchmarks, Raon-SpeechChat shows its clearest strengths on the turn-taking and interruption-sensitive behaviors covered by FDB v1.0, and remains competitive across the broader full-duplex evaluation suite. We open-source all model checkpoints, the training and inference pipeline, and an interactive demo.

📄 PDF Abstract BibTeX arXiv:2605.23912

Code (0)

등록된 구현이 없습니다.

Tasks

Knowledge DistillationQuestion Answering

Similar Papers 제목 키워드 기반

Decoding Imagined Speech using Wavelet Features and Deep Neural Networks

2020-03-19 · Jerrin Thomas Panachakel, A. G. Ramakrishnan

This paper proposes a novel approach that uses deep neural networks for classifying imagined speech, significantly increasing the classification accuracy. The proposed approach employs only the EEG channels over specific…

ClassificationEEGElectroencephalogram (EEG)General Classification

Technical Report on classification of literature related to children speech disorder

2025-05-20 · Ziang Wang, Amir Aryani

This technical report presents a natural language processing (NLP)-based approach for systematically classifying scientific literature on childhood speech disorders. We retrieved and filtered 4,804 relevant articles publ…

Articles

The X-LANCE Technical Report for Interspeech 2024 Speech Processing Using Discrete Speech Unit Challenge

2024-04-09 · Yiwei Guo, Chenrun Wang, Yifan Yang, Hankun Wang 외

Discrete speech tokens have been more and more popular in multiple speech processing fields, including automatic speech recognition (ASR), text-to-speech (TTS) and singing voice synthesis (SVS). In this paper, we describ…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Singing Voice Synthesisspeech-recognition+3

VibeVoice Technical Report

2025-08-26 · Zhiliang Peng, Jianwei Yu, Wenhui Wang, Yaoyao Chang 외 arxiv

This report presents VibeVoice, a novel model designed to synthesize long-form speech with multiple speakers by employing next-token diffusion, which is a unified method for modeling continuous data by autoregressively g…

Computational Efficiency

Neural Speech Synthesis for Estonian

2020-10-06 · Liisa Rätsep, Liisi Piits, Hille Pajupuu, Indrek Hein 외

This technical report describes the results of a collaboration between the NLP research group at the University of Tartu and the Institute of Estonian Language on improving neural speech synthesis for Estonian. The repor…

SentenceSpeech Synthesistext-to-speechText to Speech