paper-with-me

홈 › Papers

Anim-Director: A Large Multimodal Model Powered Agent for Controllable Animation Video Generation

2024-08-19 · Yunxin Li, Haoyuan Shi, Baotian Hu, Longyue Wang, Jiashun Zhu, Jinyi Xu, Zhen Zhao, Min Zhang

Traditional animation generation methods depend on training generative models with human-labelled data, entailing a sophisticated multi-stage pipeline that demands substantial human effort and incurs high training costs. Due to limited prompting plans, these methods typically produce brief, information-poor, and context-incoherent animations. To overcome these limitations and automate the animation process, we pioneer the introduction of large multimodal models (LMMs) as the core processor to build an autonomous animation-making agent, named Anim-Director. This agent mainly harnesses the advanced understanding and reasoning capabilities of LMMs and generative AI tools to create animated videos from concise narratives or simple instructions. Specifically, it operates in three main stages: Firstly, the Anim-Director generates a coherent storyline from user inputs, followed by a detailed director's script that encompasses settings of character profiles and interior/exterior descriptions, and context-coherent scene descriptions that include appearing characters, interiors or exteriors, and scene events. Secondly, we employ LMMs with the image generation tool to produce visual images of settings and scenes. These images are designed to maintain visual consistency across different scenes using a visual-language prompting method that combines scene descriptions and images of the appearing character and setting. Thirdly, scene images serve as the foundation for producing animated videos, with LMMs generating prompts to guide this process. The whole process is notably autonomous without manual intervention, as the LMMs interact seamlessly with generative tools to generate prompts, evaluate visual quality, and select the best one to optimize the final output.

📄 PDF Abstract BibTeX arXiv:2408.09787

Code (1)

hitsz-tmg/anim-director 공식 구현 pytorch

Tasks

Image GenerationVideo Generation

Similar Papers 제목 키워드 기반

AniME: Adaptive Multi-Agent Planning for Long Animation Generation

2025-08-26 · Lisai Zhang, Baohan Xu, Siqian Yang, Mingyu Yin 외 arxiv

We present AniME, a director-oriented multi-agent system for automated long-form anime production, covering the full workflow from a story to the final video. The director agent keeps a global memory for the whole workfl…

Towards Authentic Movie Dubbing with Retrieve-Augmented Director-Actor Interaction Learning

2025-11-18 · Rui Liu, Yuan Zhao, Zhenqi Jia arxiv

The automatic movie dubbing model generates vivid speech from given scripts, replicating a speaker's timbre from a brief timbre prompt while ensuring lip-sync with the silent video. Existing approaches simulate a simplif…

AnimAgents: Coordinating Multi-Stage Animation Pre-Production with Human-Multi-Agent Collaboration

2025-11-22 · Wen-Fan Wang, Chien-Ting Lu, Jin Ping Ng, Yi-Ting Chiu 외 arxiv

Animation pre-production lays the foundation of an animated film by transforming initial concepts into a coherent blueprint across interdependent stages such as ideation, scripting, design, and storyboarding. While gener…

Image Generation

Co-Director: Agentic Generative Video Storytelling

2026-04-27 · Yale Song, Yiwen Song, Nick Losier, Nathan Hodson 외 arxiv

While diffusion models generate high-fidelity video clips, transforming them into coherent storytelling engines remains challenging. Current agentic pipelines automate this via chained modules but suffer from semantic dr…

The Script is All You Need: An Agentic Framework for Long-Horizon Dialogue-to-Cinematic Video Generation

2026-01-25 · Chenyu Mu, Xin He, Qu Yang, Wanshun Chen 외 arxiv

Recent advances in video generation have produced models capable of synthesizing stunning visual content from simple text prompts. However, these models struggle to generate long-form, coherent narratives from high-level…

Video Generation