paper-with-me

Papers

Simulating the Visual World with Artificial Intelligence: A Roadmap

2025-11-11 · Jingtong Yue, Ziqi Huang, Zhaoxi Chen, Xintao Wang, Pengfei Wan, Ziwei Liu arxiv

The landscape of video generation is shifting, from a focus on generating visually appealing clips to building virtual environments that support interaction and maintain physical plausibility. These developments point toward the emergence of video foundation models that function not only as visual generators but also as implicit world models, models that simulate the physical dynamics, agent-environment interactions, and task planning that govern real or imagined worlds. This survey provides a systematic overview of this evolution, conceptualizing modern video foundation models as the combination of two core components: an implicit world model and a video renderer. The world model encodes structured knowledge about the world, including physical laws, interaction dynamics, and agent behavior. It serves as a latent simulation engine that enables coherent visual reasoning, long-term temporal consistency, and goal-driven planning. The video renderer transforms this latent simulation into realistic visual observations, effectively producing videos as a "window" into the simulated world. We trace the progression of video generation through four generations, in which the core capabilities advance step by step, ultimately culminating in a world model, built upon a video generation model, that embodies intrinsic physical plausibility, real-time multimodal interaction, and planning capabilities spanning multiple spatiotemporal scales. For each generation, we define its core characteristics, highlight representative works, and examine their application domains such as robotics, autonomous driving, and interactive gaming. Finally, we discuss open challenges and design principles for next-generation world models, including the role of agent intelligence in shaping and evaluating these systems. An up-to-date list of related works is maintained at this link.

📄 PDF Abstract BibTeX arXiv:2511.08585

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingVisual ReasoningVideo Generation

Similar Papers 제목 키워드 기반

A 20-Year Community Roadmap for Artificial Intelligence Research in the US

2019-08-07 · Yolanda Gil, Bart Selman

Decades of research in artificial intelligence (AI) have produced formidable technologies that are providing immense benefit to industry, government, and society. AI systems can now translate across multiple languages, i…

Darwin Mobile Agent: A Roadmap for Self-Evolution

2026-05-26 · Daniel Beechey, Derek Yuen, Jianheng Liu, Dezhao Luo 외 arxiv

The goal of artificial intelligence is to create agents capable of general, adaptive behaviour in open-ended environments. Guided by the "Bitter Lesson", we argue that the most effective path toward this goal is to syste…

Reinforcement Learning

Toward Next-Generation Artificial Intelligence: Catalyzing the NeuroAI Revolution

2022-10-15 · Anthony Zador, Sean Escola, Blake Richards, Bence Ölveczky 외

Neuroscience has long been an essential driver of progress in artificial intelligence (AI). We propose that to accelerate progress in AI, we must invest in fundamental research in NeuroAI. A core component of this is the…

Prospective Artificial Intelligence Approaches for Active Cyber Defence

2021-04-20 · Neil Dhir, Henrique Hoeltgebaum, Niall Adams, Mark Briers 외

Cybercriminals are rapidly developing new malicious tools that leverage artificial intelligence (AI) to enable new classes of adaptive and stealthy attacks. New defensive methods need to be developed to counter these thr…

Causal InferencePositionreinforcement-learningReinforcement Learning (RL)

From Generative AI to Innovative AI: An Evolutionary Roadmap

2025-03-14 · Seyed Mahmoud Sajjadi Mohammadabadi

This paper explores the critical transition from Generative Artificial Intelligence (GenAI) to Innovative Artificial Intelligence (InAI). While recent advancements in GenAI have enabled systems to produce high-quality co…