paper-with-me

Papers

Generative Modelling for Controllable Audio Synthesis of Expressive Piano Performance

2020-06-16 · Hao Hao Tan, Yin-Jyun Luo, Dorien Herremans

We present a controllable neural audio synthesizer based on Gaussian Mixture Variational Autoencoders (GM-VAE), which can generate realistic piano performances in the audio domain that closely follows temporal conditions of two essential style features for piano performances: articulation and dynamics. We demonstrate how the model is able to apply fine-grained style morphing over the course of synthesizing the audio. This is based on conditions which are latent variables that can be sampled from the prior or inferred from other pieces. One of the envisioned use cases is to inspire creative and brand new interpretations for existing pieces of piano music.

📄 PDF Abstract BibTeX arXiv:2006.09833

Code (1)

gudgud96/piano-synthesis 공식 구현 pytorch

Tasks

Audio Synthesis

Similar Papers 제목 키워드 기반

Generative Adversarial Training for Text-to-Speech Synthesis Based on Raw Phonetic Input and Explicit Prosody Modelling

2023-10-14 · Tiberiu Boros, Stefan Daniel Dumitrescu, Ionut Mironica, Radu Chivereanu

We describe an end-to-end speech synthesis system that uses generative adversarial training. We train our Vocoder for raw phoneme-to-audio conversion, using explicit phonetic, pitch and duration modeling. We experiment w…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis+1

Annotation-Free MIDI-to-Audio Synthesis via Concatenative Synthesis and Generative Refinement

2024-10-22 · Osamu Take, Taketo Akama

Recent MIDI-to-audio synthesis methods using deep neural networks have successfully generated high-quality, expressive instrumental tracks. However, these methods require MIDI annotations for supervised training, limitin…

Audio SynthesisDiversity

MIDI-VALLE: Improving Expressive Piano Performance Synthesis Through Neural Codec Language Modelling

2025-07-11 · Jingjing Tang, Xin Wang, Zhe Zhang, Junichi Yamagishi 외

Generating expressive audio performances from music scores requires models to capture both instrument acoustics and human interpretation. Traditional music performance synthesis pipelines follow a two-stage approach, fir…

Audio SynthesisLanguage Modellingtext-to-speechText to Speech

STYLER: Style Factor Modeling with Rapidity and Robustness via Speech Decomposition for Expressive and Controllable Neural Text to Speech

2021-03-17 · Keon Lee, Kyumin Park, Daeyoung Kim

Previous works on neural text-to-speech (TTS) have been addressed on limited speed in training and inference time, robustness for difficult synthesis conditions, expressiveness, and controllability. Although several appr…

Speech SynthesisStyle Transfertext-to-speechText to Speech

Enhancing audio quality for expressive Neural Text-to-Speech

2021-08-13 · Abdelhamid Ezzerg, Adam Gabrys, Bartosz Putrycz, Daniel Korzekwa 외

Artificial speech synthesis has made a great leap in terms of naturalness as recent Text-to-Speech (TTS) systems are capable of producing speech with similar quality to human recordings. However, not all speaking styles …

Acoustic ModellingSpeech Synthesistext-to-speechText to Speech