paper-with-me

Papers

Effects of Convolutional Autoencoder Bottleneck Width on StarGAN-based Singing Technique Conversion

2023-08-19 · Tung-Cheng Su, Yung-Chuan Chang, Yi-Wen Liu

Singing technique conversion (STC) refers to the task of converting from one voice technique to another while leaving the original singer identity, melody, and linguistic components intact. Previous STC studies, as well as singing voice conversion research in general, have utilized convolutional autoencoders (CAEs) for conversion, but how the bottleneck width of the CAE affects the synthesis quality has not been thoroughly evaluated. To this end, we constructed a GAN-based multi-domain STC system which took advantage of the WORLD vocoder representation and the CAE architecture. We varied the bottleneck width of the CAE, and evaluated the conversion results subjectively. The model was trained on a Mandarin dataset which features four singers and four singing techniques: the chest voice, the falsetto, the raspy voice, and the whistle voice. The results show that a wider bottleneck corresponds to better articulation clarity but does not necessarily lead to higher likeness to the target technique. Among the four techniques, we also found that the whistle voice is the easiest target for conversion, while the other three techniques as a source produce more convincing conversion results than the whistle.

📄 PDF Abstract BibTeX arXiv:2308.10021

Code (0)

등록된 구현이 없습니다.

Tasks

Voice Conversion

Similar Papers 제목 키워드 기반

Walking the Tightrope: An Investigation of the Convolutional Autoencoder Bottleneck

2019-11-18 · Ilja Manakov, Markus Rohm, Volker Tresp

In this paper, we present an in-depth investigation of the convolutional autoencoder (CAE) bottleneck. Autoencoders (AE), and especially their convolutional variants, play a vital role in the current deep learning toolbo…

Outlier DetectionRepresentation LearningTransfer Learning

Beyond Linear and Overcomplete Regimes: A Mean-Field Analysis of Bottleneck Autoencoders

2026-06-05 · Santanu Das, Ramyak Bilas, Pascal Esser, Satyaki Mukherjee arxiv

Autoencoders (AEs) learn low-dimensional representations by mapping data into a latent space while minimizing reconstruction error. Despite their empirical success, theoretical understanding remains limited and largely r…

An Improved StarGAN for Emotional Voice Conversion: Enhancing Voice Quality and Data Augmentation

2021-07-18 · Xiangheng He, Junjie Chen, Georgios Rizos, Björn W. Schuller

Emotional Voice Conversion (EVC) aims to convert the emotional style of a source speech signal to a target style while preserving its content and speaker identity information. Previous emotional conversion studies do not…

Data AugmentationEmotion RecognitionGenerative Adversarial NetworkSpeech Emotion Recognition+1

Visual anomaly detection in video by variational autoencoder

2022-03-08 · Faraz Waseem, Rafael Perez Martinez, Chris Wu

Video anomalies detection is the intersection of anomaly detection and visual intelligence. It has commercial applications in surveillance, security, self-driving cars and crop monitoring. Videos can capture a variety of…

Anomaly DetectionSelf-Driving Cars

Convolutional Autoencoder-Based Phase Shift Feedback Compression for Intelligent Reflecting Surface-Assisted Wireless Systems

2021-10-24 · Xianhua Yu, Dong Li, Yongjun Xu, Ying-Chang Liang

In recent years, intelligent reflecting surface (IRS) has emerged as a promising technology for 6G due to its potential/ability to significantly enhance energy- and spectrum-efficiency. To this end, it is crucial to adju…

Quantization