paper-with-me

Papers

RIVET: Robust Idempotent Voice Attribute Editing

2026-06-17 · Dareen Alharthi, Bhuvan Koduru, Rita Singh, Bhiksha Raj arxiv

Voice attribute editing models modify characteristics such as age and gender while preserving speaker identity. In large-scale speech datasets, however, attribute annotations are often noisy or inconsistent, which can cause conditional generative models to produce unstable edits. In this work, we show that idempotency provides an effective mechanism for improving robustness to noisy labels. An idempotent operator is one for which repeated application does not change the result, i.e., f(f(x)) = f(x). Enforcing this property acts as an implicit regularizer that reduces sensitivity to mislabeled examples. We introduce RIVET, a training framework that incorporates an idempotency objective to improve robustness to label noise. We evaluate RIVET under controlled label noise and on the GLOBE dataset with naturally noisy annotations. RIVET improves editing success and better preserves speaker identity than standard training, showing that idempotency improves robustness in voice editing models.

📄 PDF Abstract BibTeX arXiv:2606.19629

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Voice Attribute Editing with Text Prompt

2024-04-13 · Zhengyan Sheng, Yang Ai, Li-Juan Liu, Jia Pan 외

Despite recent advancements in speech generation with text prompt providing control over speech style, voice attributes in synthesized speech remain elusive and challenging to control. This paper introduces a novel task:…

Attribute

DriveTrack: A Benchmark for Long-Range Point Tracking in Real-World Videos

2023-12-15 · CVPR 2024 1 · Arjun Balasingam, Joseph Chandler, Chenning Li, Zhoutong Zhang 외

This paper presents DriveTrack, a new benchmark and data generation framework for long-range keypoint tracking in real-world videos. DriveTrack is motivated by the observation that the accuracy of state-of-the-art tracke…

Autonomous DrivingPoint Tracking

LILAC: An Idempotent Neural Speech Codec

2026-08-06 · June Young Yi, Dongwook Lee, Jiheum Yeom, Sungroh Yoon arxiv

Neural Audio Codecs are widely adopted in speech generation and editing. However, existing neural audio codecs are not idempotent: across the paper's twelve baseline systems, every configuration tested rewrites, on avera…

VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing

2024-04-10 · Philip Anastassiou, Zhenyu Tang, Kainan Peng, Dongya Jia 외

We present VoiceShop, a novel speech-to-speech framework that can modify multiple attributes of speech, such as age, gender, accent, and speech style, in a single forward pass while preserving the input speaker's timbre.…

Attribute

Score-based Idempotent Distillation of Diffusion Models

2025-09-25 · Shehtab Zaman, Chengyan Liu, Kenneth Chiu arxiv

Idempotent generative networks (IGNs) are a new line of generative models based on idempotent mapping to a target manifold. IGNs support both single-and multi-step generation, allowing for a flexible trade-off between co…