paper-with-me

Papers

Marco-Voice Technical Report

2025-08-04 · Fengping Tian, Chenyang Lyu, Xuanfan Ni, Haoqin Sun, Qingjuan Li, Zhiqiang Qian, Haijun Li, Longyue Wang, Zhao Xu, Weihua Luo, Kaifu Zhang arxiv

This paper presents a multifunctional speech synthesis system that integrates voice cloning and emotion control speech synthesis within a unified framework. The goal of this work is to address longstanding challenges in achieving highly expressive, controllable, and natural speech generation that faithfully preserves speaker identity across diverse linguistic and emotional contexts. Our approach introduces an effective speaker-emotion disentanglement mechanism with in-batch contrastive learning, enabling independent manipulation of speaker identity and eemotional style, as well as rotational emotional embedding integration method for smooth emotion control. To support comprehensive training and evaluation, we construct CSEMOTIONS, a high-quality emotional speech dataset containing 10 hours of Mandarin speech from ten professional speakers across seven emotional categories. Extensive experiments demonstrate that our system, Marco-Voice, achieves substantial improvements in both objective and subjective metrics. Comprehensive evaluations and analysis were conducted, results show that MarcoVoice delivers competitive performance in terms of speech clarity and emotional richness, representing a substantial advance in the field of expressive neural speech synthesis. Our code and dataset are publicly available at https://github.com/AIDC-AI/Marco-Voice and https://huggingface.co/datasets/AIDC-AI/CSEMOTIONS respectively.

📄 PDF Abstract BibTeX arXiv:2508.02038

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningSpeech Synthesis

Similar Papers 제목 키워드 기반

VIBEVOICE-ASR Technical Report

2026-01-26 · Zhiliang Peng, Jianwei Yu, Yaoyao Chang, Zilong Wang 외 arxiv

This report presents VibeVoice-ASR, a general-purpose speech understanding framework built upon VibeVoice, designed to address the persistent challenges of context fragmentation and multi-speaker complexity in long-form …

Speaker DiarizationSpeech Recognition

VibeVoice Technical Report

2025-08-26 · Zhiliang Peng, Jianwei Yu, Wenhui Wang, Yaoyao Chang 외 arxiv

This report presents VibeVoice, a novel model designed to synthesize long-form speech with multiple speakers by employing next-token diffusion, which is a unified method for modeling continuous data by autoregressively g…

Computational Efficiency

The tale of two MS MARCO -- and their unfair comparisons

2023-04-25 · Carlos Lassance, Stéphane Clinchant

The MS MARCO-passage dataset has been the main large-scale dataset open to the IR community and it has fostered successfully the development of novel neural retrieval models over the years. But, it turns out that two dif…

RetrievalVocal Bursts Valence Prediction

Creation and Detection of German Voice Deepfakes

2021-08-02 · Vanessa Barnekow, Dominik Binder, Niclas Kromrey, Pascal Munaretto 외

Synthesizing voice with the help of machine learning techniques has made rapid progress over the last years [1] and first high profile fraud cases have been recently reported [2]. Given the current increase in using conf…

AvgBIG-bench Machine Learning

VibeVoice-ASR-BitNet Technical Report

2026-07-23 · Songchen Xu, Ting Song, Shaohan Huang, Zhiliang Peng 외 arxiv

We present VibeVoice-ASR-BitNet, a compressed variant of VibeVoice-ASR optimized for real-time inference on edge CPUs. We apply heterogeneous quantization tailored to the computational characteristics of each stage: the …