paper-with-me

Papers

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion

2025-07-19 · Yu Zhang, Baotong Tian, Zhiyao Duan arxiv

Zero-shot online voice conversion (VC) holds significant promise for real-time communications and entertainment. However, current VC models struggle to preserve semantic fidelity under real-time constraints, deliver natural-sounding conversions, and adapt effectively to unseen speaker characteristics. To address these challenges, we introduce Conan, a chunkwise online zero-shot voice conversion model that preserves the content of the source while matching the voice timbre and styles of reference speech. Conan comprises three core components: 1) a Stream Content Extractor that leverages Emformer for low-latency streaming content encoding; 2) an Adaptive Style Encoder that extracts fine-grained stylistic features from reference speech for enhanced style adaptation; 3) a Causal Shuffle Vocoder that implements a fully causal HiFiGAN using a pixel-shuffle mechanism. Experimental evaluations demonstrate that Conan outperforms baseline models in subjective and objective metrics. Audio samples can be found at https://aaronz345.github.io/ConanDemo.

📄 PDF Abstract BibTeX arXiv:2507.14534

Code (0)

등록된 구현이 없습니다.

Tasks

Voice Conversion

Similar Papers 제목 키워드 기반

Continuous Entailment Patterns for Lexical Inference in Context

2021-09-08 · EMNLP 2021 11 · Martin Schmitt, Hinrich Schütze

Combining a pretrained language model (PLM) with textual patterns has been shown to help in both zero- and few-shot settings. For zero-shot performance, it makes sense to design patterns that closely resemble the text se…

Few-Shot NLILanguage ModelingLexical EntailmentNatural Language Understanding

Monotonic Chunkwise Attention

2017-12-14 · Chung-Cheng Chiu, Colin Raffel

Sequence-to-sequence models with soft attention have been successfully applied to a wide variety of problems, but their decoding process incurs a quadratic time and space cost and is inapplicable to real-time sequence tr…

Document Summarizationspeech-recognitionSpeech Recognition

Monotonic Chunkwise Attention

2018-01-01 · ICLR 2018 1 · Chung-Cheng Chiu*, Colin Raffel*

Sequence-to-sequence models with soft attention have been successfully applied to a wide variety of problems, but their decoding process incurs a quadratic time and space cost and is inapplicable to real-time sequence tr…

Document Summarizationspeech-recognitionSpeech Recognition

X-VC: Zero-shot Streaming Voice Conversion in Codec Space

2026-04-14 · Qixi Zheng, Yuxiang Zhao, Tianrui Wang, Wenxi Chen 외 arxiv

Zero-shot voice conversion (VC) aims to convert a source utterance into the voice of an unseen target speaker while preserving its linguistic content. Although recent systems have improved conversion quality, building ze…

Voice Conversion

FC-CONAN: An Exhaustively Paired Dataset for Robust Evaluation of Retrieval Systems

2026-01-04 · Juan Junqueras, Florian Boudin, May-Myo Zin, Ha-Thanh Nguyen 외 arxiv

Hate speech (HS) is a critical issue in online discourse, and one promising strategy to counter it is through the use of counter-narratives (CNs). Datasets linking HS with CNs are essential for advancing counterspeech re…