paper-with-me

홈 › Papers

Audio Editing with Non-Rigid Text Prompts

2023-10-19 · Francesco Paissan, Luca Della Libera, Zhepei Wang, Mirco Ravanelli, Paris Smaragdis, Cem Subakan

In this paper, we explore audio-editing with non-rigid text edits. We show that the proposed editing pipeline is able to create audio edits that remain faithful to the input audio. We explore text prompts that perform addition, style transfer, and in-painting. We quantitatively and qualitatively show that the edits are able to obtain results which outperform Audio-LDM, a recently released text-prompted audio generation model. Qualitative inspection of the results points out that the edits given by our approach remain more faithful to the input audio in terms of keeping the original onsets and offsets of the audio events.

📄 PDF Abstract BibTeX arXiv:2310.12858

Code (0)

등록된 구현이 없습니다.

Tasks

Audio GenerationStyle Transfer

Similar Papers 제목 키워드 기반

Unified Diffusion-Based Rigid and Non-Rigid Editing with Text and Image Guidance

2024-01-04 · Jiacheng Wang, Ping Liu, Wei Xu

Existing text-to-image editing methods tend to excel either in rigid or non-rigid editing but encounter challenges when combining both, resulting in misaligned outputs with the provided text prompts. In addition, integra…

Appearance Transfer

Audio-Guided Visual Editing with Complex Multi-Modal Prompts

2025-08-28 · Hyeonyu Kim, Seokhoon Jeong, Seonghee Han, Chanhyuk Choi 외 arxiv

Visual editing with diffusion models has made significant progress but often struggles with complex scenarios that textual guidance alone could not adequately describe, highlighting the need for additional non-text editi…

MEDIC: Zero-shot Music Editing with Disentangled Inversion Control

2024-07-18 · Huadai Liu, Jialei Wang, Xiangtai Li, Rongjie Huang 외

Text-guided diffusion models make a paradigm shift in audio generation, facilitating the adaptability of source audio to conform to specific textual prompts. Recent works introduce inversion techniques, like DDIM inversi…

Audio Generation

In-Context Prompt Editing For Conditional Audio Generation

2023-11-01 · Ernie Chang, Pin-Jie Lin, Yang Li, Sidd Srinivasan 외

Distributional shift is a central challenge in the deployment of machine learning models as they can be ill-equipped for real-world data. This is particularly evident in text-to-audio generation where the encoded represe…

Audio GenerationRetrieval

FlexiEdit: Frequency-Aware Latent Refinement for Enhanced Non-Rigid Editing

2024-07-25 · Gwanhyeong Koo, Sunjae Yoon, Ji Woo Hong, Chang D. Yoo

Current image editing methods primarily utilize DDIM Inversion, employing a two-branch diffusion approach to preserve the attributes and layout of the original image. However, these methods encounter challenges with non-…

Text-based Image Editing