paper-with-me

홈 › Papers

MAGIC: Transition-Aware Generation of Navigable Multi-Scene Game Worlds with Large Language Models

2026-07-13 · Tsz Hei Fan, Choi Wing Fung, Yuxuan Wan, Shuqing Li, Michael R. Lyu hf

Multi-scene navigation (clearing an objective in one bounded space and then crossing a portal into the next) is a defining feature of contemporary 3D games, but authoring it is laborious: every portal must have consistent endpoints on both sides, each interior must remain navigable once it is furnished, and the resulting connectivity must be kept consistent across many files. Recent large language model (LLM) and multimodal LLM (MLLM) scene generators have made single-interior synthesis dramatically cheaper, yet they produce one scene at a time and cannot, by naive repetition, yield a connected multi-scene world. We identify three obstacles that single-scene methods leave unsolved: cross-scene consistency, in-scene navigability, and the evaluation of whether a transition actually works. We present MAGIC, a prompt-to-project system that addresses all three. MAGIC is a four-stage pipeline that turns a single natural-language prompt into a runnable multi-scene game project: it plans a shared transition-aware intermediate representation, specifies each scene while enforcing portal reachability with a flood-fill validator, generates the scenes together with their transition scripts, and combines them into one project. Because existing single-scene fidelity metrics never execute a transition, we further introduce a transition-focused evaluation agent that runs each transition in play. On a new benchmark of 100 multi-scene cases, MAGIC produces an executable project for every case and reaches 0.99 precision, 0.95 recall, and 0.96 F1 on end-to-end transition identification; stage by stage, it recovers more ground-truth portals and yields markedly more navigable layouts than an LLM baseline and Holodeck. Our code is available at https://github.com/sereneee1201/MAGIC/.

📄 PDF Abstract BibTeX arXiv:2607.11594

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MAGIC-Talk: Motion-aware Audio-Driven Talking Face Generation with Customizable Identity Control

2025-10-26 · Fatemeh Nazarieh, Zhenhua Feng, Diptesh Kanojia, Muhammad Awais 외 arxiv

Audio-driven talking face generation has gained significant attention for applications in digital media and virtual avatars. While recent methods improve audio-lip synchronization, they often struggle with temporal consi…

Talking Face GenerationVideo Generation

MagicGeo: Training-Free Text-Guided Geometric Diagram Generation

2025-02-19 · Junxiao Wang, Ting Zhang, Heng Yu, Jingdong Wang 외

Geometric diagrams are critical in conveying mathematical and scientific concepts, yet traditional diagram generation methods are often manual and resource-intensive. While text-to-image generation has made strides in ph…

Image GenerationText to Image GenerationText-to-Image Generation

MagicAvatar: Multimodal Avatar Generation and Animation

2023-08-28 · Jianfeng Zhang, Hanshu Yan, Zhongcong Xu, Jiashi Feng 외

This report presents MagicAvatar, a framework for multimodal video generation and animation of human avatars. Unlike most existing methods that generate avatar-centric videos directly from multimodal inputs (e.g., text p…

Video Generation

MagicProp: Diffusion-based Video Editing via Motion-aware Appearance Propagation

2023-09-02 · Hanshu Yan, Jun Hao Liew, Long Mai, Shanchuan Lin 외

This paper addresses the issue of modifying the visual appearance of videos while preserving their motion. A novel framework, named MagicProp, is proposed, which disentangles the video editing process into two stages: ap…

Video Editing

MagicScroll: Nontypical Aspect-Ratio Image Generation for Visual Storytelling via Multi-Layered Semantic-Aware Denoising

2023-12-18 · Bingyuan Wang, Hengyu Meng, Zeyu Cai, Lanjiong Li 외

Visual storytelling often uses nontypical aspect-ratio images like scroll paintings, comic strips, and panoramas to create an expressive and compelling narrative. While generative AI has achieved great success and shown …

DenoisingImage GenerationVisual Storytelling