paper-with-me

Papers

MoReGen: Multi-Agent Motion-Reasoning Engine for Code-based Text-to-Video Synthesis

2025-12-03 · Xiangyu Bai, He Liang, Bishoy Galoaa, Utsav Nandi, Shayda Moezzi, Yuhang He, Sarah Ostadabbas arxiv

While text-to-video (T2V) generation has achieved remarkable progress in photorealism, generating intent-aligned videos that faithfully obey physics principles remains a core challenge. In this work, we systematically study Newtonian motion-controlled text-to-video generation and evaluation, emphasizing physical precision and motion coherence. We introduce MoReGen, a motion-aware, physics-grounded T2V framework that integrates multi-agent LLMs, physics simulators, and renderers to generate reproducible, physically accurate videos from text prompts in the code domain. To quantitatively assess physical validity, we propose object-trajectory correspondence as a direct evaluation metric and present MoReSet, a benchmark of 1,275 human-annotated videos spanning nine classes of Newtonian phenomena with scene descriptions, spatiotemporal relations, and ground-truth trajectories. Using MoReSet, we conduct experiments on existing T2V models, evaluating their physical validity through both our MoRe metrics and existing physics-based evaluators. Our results reveal that state-of-the-art models struggle to maintain physical validity, while MoReGen establishes a principled direction toward physically coherent video synthesis.

📄 PDF Abstract BibTeX arXiv:2512.04221

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Video Generation

Similar Papers 제목 키워드 기반

Levels of AI Agents: from Rules to Large Language Models

2024-03-06 · Yu Huang

AI agents are defined as artificial entities to perceive the environment, make decisions and take actions. Inspired by the 6 levels of autonomous driving by Society of Automotive Engineers, the AI agents are also categor…

Autonomous DrivingDecision Making

Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark

2025-08-23 · Jinpeng Hu, Hongchang Shi, Chongyuan Dai, Zhuo Li 외 arxiv

Multimodal large language models (MLLMs) have been widely applied across various fields due to their powerful perceptual and reasoning capabilities. In the realm of psychology, these models hold promise for a deeper unde…

Emotion Recognition

EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization

2026-07-23 · Lihuang Fang, Yuchen Zou, kebin Jin, Jinghui Qin arxiv

Multimodal large language models (MLLMs) have achieved impressive performance in multimodal emotion recognition (MER) tasks and lifted MER to a new level that is complex emotion understanding with advanced video understa…

Multimodal Emotion RecognitionReinforcement Learning

Cutscene Agent: An LLM Agent Framework for Automated 3D Cutscene Generation

2026-04-28 · Lanshan He, Haozhou Pang, Qi Gan, Xin Shen 외 arxiv

Cutscenes are carefully choreographed cinematic sequences embedded in video games and interactive media, serving as the primary vehicle for narrative delivery, character development, and emotional engagement. Producing c…

Visual Reasoning

Unifying Language Agent Algorithms with Graph-based Orchestration Engine for Reproducible Agent Research

2025-05-30 · Qianqian Zhang, Jiajia Liao, Heting Ying, Yibo Ma 외

Language agents powered by large language models (LLMs) have demonstrated remarkable capabilities in understanding, reasoning, and executing complex tasks. However, developing robust agents presents significant challenge…

Mathematical Reasoning