paper-with-me

Papers

TransAnimate: Taming Layer Diffusion to Generate RGBA Video

2025-03-23 · Xuewei Chen, Zhimin Chen, Yiren Song

Text-to-video generative models have made remarkable advancements in recent years. However, generating RGBA videos with alpha channels for transparency and visual effects remains a significant challenge due to the scarcity of suitable datasets and the complexity of adapting existing models for this purpose. To address these limitations, we present TransAnimate, an innovative framework that integrates RGBA image generation techniques with video generation modules, enabling the creation of dynamic and transparent videos. TransAnimate efficiently leverages pre-trained text-to-transparent image model weights and combines them with temporal models and controllability plugins trained on RGB videos, adapting them for controllable RGBA video generation tasks. Additionally, we introduce an interactive motion-guided control mechanism, where directional arrows define movement and colors adjust scaling, offering precise and intuitive control for designing game effects. To further alleviate data scarcity, we have developed a pipeline for creating an RGBA video dataset, incorporating high-quality game effect videos, extracted foreground objects, and synthetic transparent videos. Comprehensive experiments demonstrate that TransAnimate generates high-quality RGBA videos, establishing it as a practical and effective tool for applications in gaming and visual effects.

📄 PDF Abstract BibTeX arXiv:2503.17934

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationVideo Generation

Similar Papers 제목 키워드 기반

Generating Compositional Scenes via Text-to-image RGBA Instance Generation

2024-11-16 · Alessandro Fontanella, Petru-Daniel Tudosiu, Yongxin Yang, Shifeng Zhang 외

Text-to-image diffusion generative models can generate high quality images at the cost of tedious prompt engineering. Controllability can be improved by introducing layout conditioning, however existing methods lack layo…

ObjectPrompt Engineering

LaDe: Unified Multi-Layered Graphic Media Generation and Decomposition

2026-03-18 · Vlad-Constantin Lungu-Stan, Ionut Mironica, Mariana-Iuliana Georgescu arxiv

Media design layer generation enables the creation of fully editable, layered design documents such as posters, flyers, and logos using only natural language prompts. Existing methods either restrict outputs to a fixed n…

Text-to-Image Generation

AlphaVAE: Unified End-to-End RGBA Image Reconstruction and Generation with Alpha-Aware Representation Learning

2025-07-12 · Zile Wang, Hao Yu, Jiabo Zhan, Chun Yuan arxiv

Recent advances in latent diffusion models have achieved remarkable results in high-fidelity RGB image synthesis by leveraging pretrained VAEs to compress and reconstruct pixel data at low computational cost. However, th…

Representation LearningImage ReconstructionImage Generation

UniWorld-Design: From Pixel Generation to Layer-Native Design

2026-08-04 · Zongjian Li, Zhiyuan Yan, Chenxu Bai, Chen Chen 외 hf

We introduce UniWorld-Design, a framework that redefines image generation from flat pixel synthesis to structured visual composition, with semantic RGBA layers as the atomic units of generation, understanding, and editin…

Image Generation

Explicit Layer Modeling for Video Object Insertion and Layer Decomposition

2026-07-28 · Kyujin Han, Seungjoo Shin, Sunghyun Cho arxiv

Most video editing systems still lack explicit layered video representations, limiting their ability to perform realistic compositing, object reuse, and consistent manipulation. This limitation is especially pronounced i…