paper-with-me

Papers

TimbreCLIP: Connecting Timbre to Text and Images

2022-11-21 · Nicolas Jonason, Bob L. T. Sturm

We present work in progress on TimbreCLIP, an audio-text cross modal embedding trained on single instrument notes. We evaluate the models with a cross-modal retrieval task on synth patches. Finally, we demonstrate the application of TimbreCLIP on two tasks: text-driven audio equalization and timbre to image generation.

📄 PDF Abstract BibTeX arXiv:2211.11225

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Modal RetrievalImage GenerationRetrieval

Similar Papers 제목 키워드 기반

OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model

2026-02-12 · Maomao Li, Zhen Li, Kaipeng Zhang, Guosheng Yin 외 arxiv

Existing mainstream video customization methods focus on generating identity-consistent videos based on given reference images and textual prompts. Benefiting from the rapid advancement of joint audio-video generation, t…

Contrastive LearningVideo Generation

GenerTTS: Pronunciation Disentanglement for Timbre and Style Generalization in Cross-Lingual Text-to-Speech

2023-06-27 · Yahuan Cong, Haoyu Zhang, Haopeng Lin, Shichao Liu 외

Cross-lingual timbre and style generalizable text-to-speech (TTS) aims to synthesize speech with a specific reference timbre or style that is never trained in the target language. It encounters the following challenges: …

DisentanglementStyle Generalizationtext-to-speechText to Speech

Zero-shot Voice Conversion with Diffusion Transformers

2024-11-15 · Songting Liu

Zero-shot voice conversion aims to transform a source speech utterance to match the timbre of a reference speech from an unseen speaker. Traditional approaches struggle with timbre leakage, insufficient timbre representa…

In-Context LearningVoice Conversion

Real-time Timbre Remapping with Differentiable DSP

2024-07-05 · Jordie Shier, Charalampos Saitis, Andrew Robertson, Andrew McPherson

Timbre is a primary mode of expression in diverse musical contexts. However, prevalent audio-driven synthesis methods predominantly rely on pitch and loudness envelopes, effectively flattening timbral expression from the…

CTEFM-VC: Zero-Shot Voice Conversion Based on Content-Aware Timbre Ensemble Modeling and Flow Matching

2024-11-04 · Yu Pan, Yuguang Yang, Jixun Yao, Jianhao Ye 외

Zero-shot voice conversion (VC) aims to transform the timbre of a source speaker into any previously unseen target speaker, while preserving the original linguistic content. Despite notable progress, attaining a degree o…

Speaker VerificationVoice Conversion