paper-with-me

Papers

Towards Lightweight and Stable Zero-shot TTS with Self-distilled Representation Disentanglement

2025-01-15 · Qianniu Chen, Xiaoyang Hao, Bowen Li, Yue Liu, Li Lu

Zero-shot Text-To-Speech (TTS) synthesis shows great promise for personalized voice customization through voice cloning. However, current methods for achieving zero-shot TTS heavily rely on large model scales and extensive training datasets to ensure satisfactory performance and generalizability across various speakers. This raises concerns regarding both deployment costs and data security. In this paper, we present a lightweight and stable zero-shot TTS system. We introduce a novel TTS architecture designed to effectively model linguistic content and various speaker attributes from source speech and prompt speech, respectively. Furthermore, we present a two-stage self-distillation framework that constructs parallel data pairs for effectively disentangling linguistic content and speakers from the perspective of training data. Extensive experiments show that our system exhibits excellent performance and superior stability on the zero-shot TTS tasks. Moreover, it shows markedly superior computational efficiency, with RTFs of 0.13 and 0.012 on the CPU and GPU, respectively.

📄 PDF Abstract BibTeX arXiv:2501.08566

Code (0)

등록된 구현이 없습니다.

Tasks

Computational EfficiencyCPUDisentanglementGPUtext-to-speechText to SpeechVoice Cloning

Similar Papers 제목 키워드 기반

Image-to-Lidar Relational Distillation for Autonomous Driving Data

2024-09-01 · Anas Mahmoud, Ali Harakeh, Steven Waslander

Pre-trained on extensive and diverse multi-modal datasets, 2D foundation models excel at addressing 2D tasks with little or no downstream supervision, owing to their robust representations. The emergence of 2D-to-3D dist…

3D Semantic SegmentationAutonomous DrivingFew-shot 3D semantic segmentationSemantic Segmentation+2

Spectrally Distilled Representations Aligned with Instruction-Augmented LLMs for Satellite Imagery

2026-02-26 · Minh Kha Do, Wei Xiang, Kang Han, Di Wu 외 arxiv

Vision-language foundation models (VLFMs) promise zero-shot and retrieval understanding for Earth observation. While operational satellite systems often lack full multi-spectral coverage, making RGB-only inference highly…

Zero-Shot Self-Consistency Learning for Seismic Irregular Spatial Sampling Reconstruction

2024-11-01 · Junheng Peng, Yingtian Liu, Mingwei Wang, Yong Li 외

Seismic exploration is currently the most important method for understanding subsurface structures. However, due to surface conditions, seismic receivers may not be uniformly distributed along the measurement line, makin…

ZSG-IAD: A Multimodal Framework for Zero-Shot Grounded Industrial Anomaly Detection

2026-04-20 · Qiuhui Chen, Jiaxiang Song, Shuai Tan, Weimin Zhong arxiv

Deep learning-based industrial anomaly detectors often behave as black boxes, making it hard to justify decisions with physically meaningful defect evidence. We propose ZSG-IAD, a multimodal vision-language framework for…

Anomaly DetectionPoint Clouds

ZipDepth: Bringing Lightweight Zero-Shot Monocular Depth Anywhere, on Any Device

2026-07-09 · Fabio Tosi, Luca Bartolomei, Matteo Poggi, Stefano Mattoccia arxiv

Monocular depth estimation has seen remarkable progress through foundation models achieving robust zero-shot generalization, yet their computational demands place them far beyond the reach of embedded and mobile platform…

Monocular Depth EstimationZero-shot GeneralizationKnowledge Distillation