paper-with-me

홈 › Papers

AnimeAgent: Is the Multi-Agent via Image-to-Video models a Good Disney Storytelling Artist?

2026-02-24 · Hailong Yan, Shice Liu, Tao Wang, Xiangtao Zhang, Yijie Zhong, Jinwei Chen, Le Zhang, Bo Li arxiv

Custom Storyboard Generation (CSG) aims to produce high-quality, multi-character consistent storytelling. Current approaches based on static diffusion models, whether used in a one-shot manner or within multi-agent frameworks, face three key limitations: (1) Static models lack dynamic expressiveness and often resort to "copy-paste" pattern. (2) One-shot inference cannot iteratively correct missing attributes or poor prompt adherence. (3) Multi-agents rely on non-robust evaluators, ill-suited for assessing stylized, non-realistic animation. To address these, we propose AnimeAgent, the first Image-to-Video (I2V)-based multi-agent framework for CSG. Inspired by Disney's "Combination of Straight Ahead and Pose to Pose" workflow, AnimeAgent leverages I2V's implicit motion prior to enhance consistency and expressiveness, while a mixed subjective-objective reviewer enables reliable iterative refinement. We also collect a human-annotated CSG benchmark with ground-truth. Experiments show AnimeAgent achieves SOTA performance in consistency, prompt fidelity, and stylization.

📄 PDF Abstract BibTeX arXiv:2602.20664

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Multi-Agent Game Generation and Evaluation via Audio-Visual Recordings

2025-08-01 · Alexia Jolicoeur-Martineau arxiv

While AI excels at generating text, audio, images, and videos, creating interactive audio-visual content such as video games remains challenging. Current LLMs can generate JavaScript games and animations, but lack automa…

Visual Summarization of Lecture Video Segments for Enhanced Navigation

2020-06-03 · Mohammad Rajiur Rahman, Jaspal Subhlok, Shishir Shah

Lecture videos are an increasingly important learning resource for higher education. However, the challenge of quickly finding the content of interest in a lecture video is an important limitation of this format. This pa…

Management

VELOCITI: Benchmarking Video-Language Compositional Reasoning with Strict Entailment

2024-06-16 · CVPR 2025 1 · Darshana Saravanan, Varun Gupta, Darshan Singh, Zeeshan Khan 외

A fundamental aspect of compositional reasoning in a video is associating people and their actions across time. Recent years have seen great progress in general-purpose vision or video models and a move towards long-vide…

Action UnderstandingBenchmarkingMultiple-choiceVideo Understanding

Unsupervised Video Object Segmentation for Deep Reinforcement Learning

2018-05-20 · NeurIPS 2018 12 · Vik Goel, Jameson Weng, Pascal Poupart

We present a new technique for deep reinforcement learning that automatically detects moving objects and uses the relevant information for action selection. The detection of moving objects is done in an unsupervised way …

Atari GamesDecision MakingDeep Reinforcement LearningObject+7

VideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web Tasks

2024-10-24 · Lawrence Jang, Yinheng Li, Charles Ding, Justin Lin 외

Videos are often used to learn or extract the necessary information to complete tasks in ways different than what text and static imagery alone can provide. However, many existing agent benchmarks neglect long-context vi…

Video Understanding