paper-with-me

홈 › Papers

MAViS: A Multi-Agent Framework for Long-Sequence Video Storytelling

2025-08-11 · Qian Wang, Ziqi Huang, Ruoxi Jia, Paul Debevec, Ning Yu arxiv

Despite recent advances, long-sequence video generation frameworks still suffer from significant limitations: poor assistive capability, suboptimal visual quality, and limited expressiveness. To mitigate these limitations, we propose MAViS, a multi-agent collaborative framework designed to assist in long-sequence video storytelling by efficiently translating ideas into visual narratives. MAViS orchestrates specialized agents across multiple stages, including script writing, shot designing, character modeling, keyframe generation, video animation, and audio generation. In each stage, agents operate under the 3E Principle -- Explore, Examine, and Enhance -- to ensure the completeness of intermediate outputs. Considering the capability limitations of current generative models, we propose the Script Writing Guidelines to optimize compatibility between scripts and generative tools. Experimental results demonstrate that MAViS achieves state-of-the-art performance in assistive capability, visual quality, and video expressiveness. Its modular framework further enables scalability with diverse generative models and tools. With just a brief idea description, MAViS enables users to rapidly explore diverse visual storytelling and creative directions for sequential video generation by efficiently producing high-quality, complete long-sequence videos. To the best of our knowledge, MAViS is the only framework that provides multimodal design output -- videos with narratives and background music.

📄 PDF Abstract BibTeX arXiv:2508.08487

Code (0)

등록된 구현이 없습니다.

Tasks

Visual StorytellingAudio GenerationVideo Generation

Similar Papers 제목 키워드 기반

MAVIS: Multi-Agent Video Retrieval via Structured Video Understanding

2026-06-08 · Jie Zhang, Qilang Ye, Hao Zhou, Haochen Liang 외 arxiv

The dominant paradigm in video retrieval relies on embedding-based full-corpus scanning, which suffers from inherent computational inefficiency and the semantic asymmetry between information-dense videos and sparse textu…

Video Retrieval

QMAVIS: Long Video-Audio Understanding using Fusion of Large Multimodal Models

2026-01-10 · Zixing Lin, Jiale Wang, Gee Wah Ng, Lee Onn Mak 외 arxiv

Large Multimodal Models (LMMs) for video-audio understanding have traditionally been evaluated only on shorter videos of a few minutes long. In this paper, we introduce QMAVIS (Q Team-Multimodal Audio Video Intelligent S…

Speech Recognition

MAviS: A Multimodal Conversational Assistant For Avian Species

2026-03-07 · Yevheniia Kryklyvets, Mohammed Irfan Kurpath, Sahal Shaji Mullappilly, Jinxing Zhou 외 arxiv

Fine-grained understanding and species-specific multimodal question answering are vital for advancing biodiversity conservation and ecological monitoring. However, existing multimodal large language models face challenge…

Question Answering

MAVIS: Multi-Objective Alignment via Inference-Time Value-Guided Selection

2025-08-19 · Jeremy Carleton, Debajoy Mukherjee, Srinivas Shakkottai, Dileep Kalathil arxiv

Large Language Models (LLMs) are increasingly deployed across diverse applications that demand balancing multiple, often conflicting, objectives -- such as helpfulness, harmlessness, or humor. Many traditional methods fo…

MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine

2024-07-11 · Renrui Zhang, Xinyu Wei, Dongzhi Jiang, Ziyu Guo 외

The mathematical capabilities of Multi-modal Large Language Models (MLLMs) remain under-explored with three areas to be improved: visual encoding of math diagrams, diagram-language alignment, and chain-of-thought (CoT) r…

Contrastive LearningLanguage ModellingLarge Language ModelMath+2