paper-with-me

Papers

RenderBox: Expressive Performance Rendering with Text Control

2025-02-11 · huan zhang, Akira Maezawa, Simon Dixon

Expressive music performance rendering involves interpreting symbolic scores with variations in timing, dynamics, articulation, and instrument-specific techniques, resulting in performances that capture musical can emotional intent. We introduce RenderBox, a unified framework for text-and-score controlled audio performance generation across multiple instruments, applying coarse-level controls through natural language descriptions and granular-level controls using music scores. Based on a diffusion transformer architecture and cross-attention joint conditioning, we propose a curriculum-based paradigm that trains from plain synthesis to expressive performance, gradually incorporating controllable factors such as speed, mistakes, and style diversity. RenderBox achieves high performance compared to baseline models across key metrics such as FAD and CLAP, and also tempo and pitch accuracy under different prompting tasks. Subjective evaluation further demonstrates that RenderBox is able to generate controllable expressive performances that sound natural and musically engaging, aligning well with prompts and intent.

📄 PDF Abstract BibTeX arXiv:2502.07711

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityFADMusic Performance Rendering

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

ScorePerformer: Expressive Piano Performance Rendering With Fine-Grained Control

2023-11-04 · ISMIR 2023 11 · Ilya Borovik, Vladimir Viro

We present ScorePerformer, an encoder-decoder transformer with hierarchical style encoding heads for controllable rendering of expressive piano music performances. We design a tokenized representation of symbolic score a…

DecoderMusic Performance Rendering

PianoKontext: Expressive Performance Rendering from Deadpan Context

2026-06-10 · Dmitrii Gavrilev arxiv

Expressive performance rendering (EPR) aims to generate realistic performances constrained on sequences of notes. However, flow matching audio editing models manipulate only synchronized music samples of the same duratio…

TexGS-VolVis: Expressive Scene Editing for Volume Visualization via Textured Gaussian Splatting

2025-07-18 · Kaiyuan Tang, Kuangshi Ai, Jun Han, Chaoli Wang arxiv

Advancements in volume visualization (VolVis) focus on extracting insights from 3D volumetric data by generating visually compelling renderings that reveal complex internal structures. Existing VolVis approaches have exp…

Style Transfer

ClipFace: Text-guided Editing of Textured 3D Morphable Models

2022-12-02 · Shivangi Aneja, Justus Thies, Angela Dai, Matthias Nießner

We propose ClipFace, a novel self-supervised approach for text-guided editing of textured 3D morphable model of faces. Specifically, we employ user-friendly language prompts to enable control of the expressions as well a…

Texture Synthesis

EVA: Expressive Virtual Avatars from Multi-view Videos

2025-05-21 · Hendrik Junkawitsch, Guoxing Sun, Heming Zhu, Christian Theobalt 외

With recent advancements in neural rendering and motion capture algorithms, remarkable progress has been made in photorealistic human avatar modeling, unlocking immense potential for applications in virtual reality, augm…

Neural Rendering