paper-with-me

Papers

MAVEN A Multi-Agent Framework for Multicultural Text-to-Video Generation

2026-05-16 · Shuowei Li, Yuming Zhao, Parth Bhalerao, Oana Ignat arxiv

Text-to-video (T2V) generation has rapidly progressed in visual fidelity, yet its ability to faithfully represent multiple cultures within a single prompt remains underexplored. We introduce MAVEN, a multi-agent prompt refinement framework designed to improve cultural fidelity in both mono-cultural and cross-cultural T2V generation. MAVEN decomposes prompts into person, action, and location dimensions, handled by specialized agents operating in parallel or sequentially. To support systematic evaluation, we contribute a new benchmark of 243 culturally grounded prompts and 972 corresponding videos, spanning three cultures (Chinese, American, Romanian), three action categories, and both mono-cultural and cross-cultural scenarios. Evaluations combining CLIP-based metrics, VLM-as-judge assessments, and videoquality measures show that multi-agent refinement, particularly parallel specialization, significantly improves cultural relevance while preserving visual quality and temporal consistency. The dataset and code are available at https://github.com/AIM-SCU/MAVEN

📄 PDF Abstract BibTeX arXiv:2605.16716

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Video Generation

Similar Papers 제목 키워드 기반

Multi-Agent Multimodal Models for Multicultural Text to Image Generation

2025-02-21 · Parth Bhalerao, Mounika Yalamarty, Brian Trinh, Oana Ignat

Large Language Models (LLMs) demonstrate impressive performance across various multimodal tasks. However, their effectiveness in cross-cultural contexts remains limited due to the predominantly Western-centric nature of …

Image GenerationText to Image GenerationText-to-Image Generation

MAVEN: Improving Generalization in Agentic Tool Calling

2026-05-29 · Omkar Ghugarkar, Vishvesh Bhat, Muhammad Ahmed Mohsin, Asad Aali arxiv

Generalization across agentic tool-calling environments remains a central challenge for reliable agentic reasoning systems. Although large language models achieve strong results on individual benchmarks, their ability to…

MAVEN: Multi-Agent Verification-Elaboration Network with In-Step Epistemic Auditing

2026-05-08 · Yinsheng Yao, Jiehao Tang, Zhaozhen Yang, Dawei Cheng arxiv

While explicit reasoning trajectories enhance model interpretability, existing paradigms often rely on monolithic chains that lack intermediate verification, allowing early errors to cascade unchecked. This lack of modul…

MAVEN: Multi-Agent Variational Exploration

2019-10-16 · NeurIPS 2019 12 · Anuj Mahajan, Tabish Rashid, Mikayel Samvelyan, Shimon Whiteson

Centralised training with decentralised execution is an important setting for cooperative deep multi-agent reinforcement learning due to communication constraints during execution and computational tractability in traini…

Multi-agent Reinforcement LearningReinforcement LearningSMACSMAC+

MAVEN: A Multi-stage Agentic Annotation Pipeline for Video Reasoning Tasks

2026-05-21 · Han Zhang, Wanting Jiang, Tomasz Kornuta, Tian Zheng 외 arxiv

Training Vision Language Models (VLMs) for video event reasoning requires high-quality structured annotations capturing not only what happened, but when, where, why, and with what consequence, at a scale manual labelling…

Domain Adaptation