paper-with-me

홈 › Papers

Resource-constrained stereo singing voice cancellation

2024-01-22 · Clara Borrelli, James Rae, Dogac Basaran, Matt Mcvicar, Mehrez Souden, Matthias Mauch

We study the problem of stereo singing voice cancellation, a subtask of music source separation, whose goal is to estimate an instrumental background from a stereo mix. We explore how to achieve performance similar to large state-of-the-art source separation networks starting from a small, efficient model for real-time speech separation. Such a model is useful when memory and compute are limited and singing voice processing has to run with limited look-ahead. In practice, this is realised by adapting an existing mono model to handle stereo input. Improvements in quality are obtained by tuning model parameters and expanding the training set. Moreover, we highlight the benefits a stereo model brings by introducing a new metric which detects attenuation inconsistencies between channels. Our approach is evaluated using objective offline metrics and a large-scale MUSHRA trial, confirming the effectiveness of our techniques in stringent listening tests.

📄 PDF Abstract BibTeX arXiv:2401.12068

Code (0)

등록된 구현이 없습니다.

Tasks

Music Source SeparationSpeech Separation

Similar Papers 제목 키워드 기반

SingFake: Singing Voice Deepfake Detection

2023-09-14 · Yongyi Zang, You Zhang, Mojtaba Heydari, Zhiyao Duan

The rise of singing voice synthesis presents critical challenges to artists and industry stakeholders over unauthorized voice usage. Unlike synthesized speech, synthesized singing voices are typically released in songs c…

DeepFake DetectionFace SwappingSinging Voice SynthesisSynthetic Speech Detection

LAPS-Diff: A Diffusion-Based Framework for Singing Voice Synthesis With Language Aware Prosody-Style Guided Learning

2025-07-07 · Sandipan Dhar, Mayank Gupta, Preeti Rao arxiv

The field of Singing Voice Synthesis (SVS) has seen significant advancements in recent years due to the rapid progress of diffusion-based approaches. However, capturing vocal style, genre-specific pitch inflections, and …

Low-Resource Cross-Domain Singing Voice Synthesis via Reduced Self-Supervised Speech Representations

2024-02-02 · Panos Kakoulidis, Nikolaos Ellinas, Georgios Vamvoukakis, Myrsini Christidou 외

In this paper, we propose a singing voice synthesis model, Karaoker-SSL, that is trained only on text and speech data as a typical multi-speaker acoustic model. It is a low-resource pipeline that does not utilize any sin…

Singing Voice Synthesis

Singing voice conversion with non-parallel data

2019-03-11 · Xin Chen, Wei Chu, Jinxi Guo, Ning Xu

Singing voice conversion is a task to convert a song sang by a source singer to the voice of a target singer. In this paper, we propose using a parallel data free, many-to-one voice conversion technique on singing voices…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

StyleSinger: Style Transfer for Out-of-Domain Singing Voice Synthesis

2023-12-17 · Yu Zhang, Rongjie Huang, RuiQi Li, Jinzheng He 외

Style transfer for out-of-domain (OOD) singing voice synthesis (SVS) focuses on generating high-quality singing voices with unseen styles (such as timbre, emotion, pronunciation, and articulation skills) derived from ref…

QuantizationSinging Voice SynthesisStyle Transfer