paper-with-me

홈 › Papers

Deep Performer: Score-to-Audio Music Performance Synthesis

2022-02-12 · Hao-Wen Dong, Cong Zhou, Taylor Berg-Kirkpatrick, Julian McAuley

Music performance synthesis aims to synthesize a musical score into a natural performance. In this paper, we borrow recent advances in text-to-speech synthesis and present the Deep Performer -- a novel system for score-to-audio music performance synthesis. Unlike speech, music often contains polyphony and long notes. Hence, we propose two new techniques for handling polyphonic inputs and providing a fine-grained conditioning in a transformer encoder-decoder model. To train our proposed system, we present a new violin dataset consisting of paired recordings and scores along with estimated alignments between them. We show that our proposed model can synthesize music with clear polyphony and harmonic structures. In a listening test, we achieve competitive quality against the baseline model, a conditional generative audio model, in terms of pitch accuracy, timbre and noise level. Moreover, our proposed model significantly outperforms the baseline on an existing piano dataset in overall quality.

📄 PDF Abstract BibTeX arXiv:2202.06034

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderSpeech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Methods 이 논문이 사용한 방법론

FAVOR+ 설명 없음
Performer Performer is a Transformer architecture which can estimate regular…

Similar Papers 제목 키워드 기반

Go witheFlow: Real-time Emotion Driven Audio Effects Modulation

2025-10-02 · Edmund Dervakos, Spyridon Kantarelis, Vassilis Lyberatos, Jason Liartis 외 arxiv

Music performance is a distinctly human activity, intrinsically linked to the performer's ability to convey, evoke, or express emotion. Machines cannot perform music in the human sense; they can produce, reproduce, execu…

ScorePerformer: Expressive Piano Performance Rendering With Fine-Grained Control

2023-11-04 · ISMIR 2023 11 · Ilya Borovik, Vladimir Viro

We present ScorePerformer, an encoder-decoder transformer with hierarchical style encoding heads for controllable rendering of expressive piano music performances. We design a tokenized representation of symbolic score a…

DecoderMusic Performance Rendering

MIDI-VALLE: Improving Expressive Piano Performance Synthesis Through Neural Codec Language Modelling

2025-07-11 · Jingjing Tang, Xin Wang, Zhe Zhang, Junichi Yamagishi 외

Generating expressive audio performances from music scores requires models to capture both instrument acoustics and human interpretation. Traditional music performance synthesis pipelines follow a two-stage approach, fir…

Audio SynthesisLanguage Modellingtext-to-speechText to Speech

Music Interpretation and Emotion Perception: A Computational and Neurophysiological Investigation

2025-05-16 · Vassilis Lyberatos, Spyridon Kantarelis, Ioanna Zioga, Christina Anagnostopoulou 외

This study investigates emotional expression and perception in music performance using computational and neurophysiological methods. The influence of different performance settings, such as repertoire, diatonic modal etu…

Emotion Recognition

Neural Music Synthesis for Flexible Timbre Control

2018-11-01 · Jong Wook Kim, Rachel Bittner, Aparna Kumar, Juan Pablo Bello

The recent success of raw audio waveform synthesis models like WaveNet motivates a new approach for music synthesis, in which the entire process --- creating audio samples from a score and instrument information --- is m…