paper-with-me

Papers

Qwen3-TTS Technical Report

2026-01-22 · Hangrui Hu, Xinfa Zhu, Ting He, Dake Guo, Bin Zhang, Xiong Wang, Zhifang Guo, Ziyue Jiang, Hongkun Hao, Zishan Guo, Xinyu Zhang, Pei Zhang, Baosong Yang, Jin Xu, Jingren Zhou, Junyang Lin arxiv

In this report, we present the Qwen3-TTS series, a family of advanced multilingual, controllable, robust, and streaming text-to-speech models. Qwen3-TTS supports state-of-the-art 3-second voice cloning and description-based control, allowing both the creation of entirely novel voices and fine-grained manipulation over the output speech. Trained on over 5 million hours of speech data spanning 10 languages, Qwen3-TTS adopts a dual-track LM architecture for real-time synthesis, coupled with two speech tokenizers: 1) Qwen-TTS-Tokenizer-25Hz is a single-codebook codec emphasizing semantic content, which offers seamlessly integration with Qwen-Audio and enables streaming waveform reconstruction via a block-wise DiT. 2) Qwen-TTS-Tokenizer-12Hz achieves extreme bitrate reduction and ultra-low-latency streaming, enabling immediate first-packet emission ($97\,\mathrm{ms}$) through its 12.5 Hz, 16-layer multi-codebook design and a lightweight causal ConvNet. Extensive experiments indicate state-of-the-art performance across diverse objective and subjective benchmark (e.g., TTS multilingual test set, InstructTTSEval, and our long speech test set). To facilitate community research and development, we release both tokenizers and models under the Apache 2.0 license.

📄 PDF Abstract BibTeX arXiv:2601.15621

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Qwen2.5-Coder Technical Report

2024-09-18 · Binyuan Hui, Jian Yang, Zeyu Cui, Jiaxi Yang 외

In this report, we introduce the Qwen2.5-Coder series, a significant upgrade from its predecessor, CodeQwen1.5. This series includes six models: Qwen2.5-Coder-(0.5B/1.5B/3B/7B/14B/32B). As a code-specific model, Qwen2.5-…

Code GenerationMathSynthetic Data Generation

QwenStyle: Content-Preserving Style Transfer with Qwen-Image-Edit

2026-01-08 · Shiwen Zhang, Haibin Huang, Chi Zhang, Xuelong Li arxiv

Content-Preserving Style transfer, given content and style references, remains challenging for Diffusion Transformers (DiTs) due to its internal entangled content and style features. In this technical report, we propose …

Continual LearningStyle Transfer

Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

2024-09-18 · An Yang, Beichen Zhang, Binyuan Hui, Bofei Gao 외

In this report, we present a series of math-specific large language models: Qwen2.5-Math and Qwen2.5-Math-Instruct-1.5B/7B/72B. The core innovation of the Qwen2.5 series lies in integrating the philosophy of self-improve…

GSM8KMathMathematical ReasoningMath Word Problem Solving+1

Qwen3-ASR Technical Report

2026-01-29 · Xian Shi, Xiong Wang, Zhifang Guo, Yongqi Wang 외 arxiv

In this report, we introduce Qwen3-ASR family, which includes two powerful all-in-one speech recognition models and a novel non-autoregressive speech forced alignment model. Qwen3-ASR-1.7B and Qwen3-ASR-0.6B are ASR mode…

Language IdentificationSpeech Recognition

Qwen2 Technical Report

2024-07-15 · An Yang, Baosong Yang, Binyuan Hui, Bo Zheng 외

This report introduces the Qwen2 series, the latest addition to our large language models and large multimodal models. We release a comprehensive suite of foundational and instruction-tuned language models, encompassing …

Arithmetic ReasoningGSM8KHumanEvalLanguage Modelling+4