paper-with-me

Papers

Transflower: probabilistic autoregressive dance generation with multimodal attention

2021-06-25 · Guillermo Valle-Pérez, Gustav Eje Henter, Jonas Beskow, André Holzapfel, Pierre-Yves Oudeyer, Simon Alexanderson

Dance requires skillful composition of complex movements that follow rhythmic, tonal and timbral features of music. Formally, generating dance conditioned on a piece of music can be expressed as a problem of modelling a high-dimensional continuous motion signal, conditioned on an audio signal. In this work we make two contributions to tackle this problem. First, we present a novel probabilistic autoregressive architecture that models the distribution over future poses with a normalizing flow conditioned on previous poses as well as music context, using a multimodal transformer encoder. Second, we introduce the currently largest 3D dance-motion dataset, obtained with a variety of motion-capture technologies, and including both professional and casual dancers. Using this dataset, we compare our new model against two baselines, via objective metrics and a user study, and show that both the ability to model a probability distribution, as well as being able to attend over a large motion and music context are necessary to produce interesting, diverse, and realistic dance that matches the music.

📄 PDF Abstract BibTeX arXiv:2106.13871

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Listen to Rhythm, Choose Movements: Autoregressive Multimodal Dance Generation via Diffusion and Mamba with Decoupled Dance Dataset

2026-01-06 · Oran Duan, Yinghua Shen, Yingzhu Lv, Luyang Jie 외 arxiv

Advances in generative models and sequence learning have greatly promoted research in dance motion generation, yet current methods still suffer from coarse semantic control and poor coherence in long sequences. In this w…

DanceMosaic: High-Fidelity Dance Generation with Multimodal Editability

2025-04-06 · Foram Niravbhai Shah, Parshwa Shah, Muhammad Usama Saleem, Ekkasit Pinyoanuntapong 외

Recent advances in dance generation have enabled automatic synthesis of 3D dance motions. However, existing methods still struggle to produce high-fidelity dance sequences that simultaneously deliver exceptional realism,…

Motion GenerationMotion Synthesis

Obliviate: Erasing Concepts from Autoregressive Image Generation Models

2026-06-26 · Hossein Shakibania, Jonas Henry Grebe, Tobias Braun, Ege Aktemur 외 arxiv

The widespread adoption of generative AI models has intensified concerns about misuse, including the creation of unsafe or disturbing imagery. To mitigate such issues, several concept erasure approaches have been propose…

Image Generation

Reducing Hallucination in Vision-Language Models via Stage-wise Preference Optimization under Distribution Shift

2026-05-13 · Qinwu Xu arxiv

Hallucination remains a fundamental challenge in vision-language models (VLMs), where autoregressive generation may produce linguistically plausible yet physically inconsistent or visually ungrounded responses due to lik…

Multimodal ReasoningSpatial ReasoningVisual Grounding

Bridging the Discrete-Continuous Gap: Unified Multimodal Generation via Coupled Manifold Discrete Absorbing Diffusion

2026-01-07 · Yuanfeng Xu, Yuhao Chen, Liang Lin, Guangrun Wang arxiv

The bifurcation of generative modeling into autoregressive approaches for discrete data (text) and diffusion approaches for continuous data (images) hinders the development of truly unified multimodal systems. While Mask…

multimodal generationImage Generation