paper-with-me

홈 › Papers

SeeU: Seeing the Unseen World via 4D Dynamics-aware Generation

2025-12-03 · Yu Yuan, Tharindu Wickremasinghe, Zeeshan Nadir, Xijun Wang, Yiheng Chi, Stanley H. Chan arxiv

Images and videos are discrete 2D projections of the 4D world (3D space + time). Most visual understanding, prediction, and generation operate directly on 2D observations, leading to suboptimal performance. We propose SeeU, a novel approach that learns the continuous 4D dynamics and generate the unseen visual contents. The principle behind SeeU is a new 2D$\to$4D$\to$2D learning framework. SeeU first reconstructs the 4D world from sparse and monocular 2D frames (2D$\to$4D). It then learns the continuous 4D dynamics on a low-rank representation and physical constraints (discrete 4D$\to$continuous 4D). Finally, SeeU rolls the world forward in time, re-projects it back to 2D at sampled times and viewpoints, and generates unseen regions based on spatial-temporal context awareness (4D$\to$2D). By modeling dynamics in 4D, SeeU achieves continuous and physically-consistent novel visual generation, demonstrating strong potentials in multiple tasks including unseen temporal generation, unseen spatial generation, and video editing. All data and code will be public at https://yuyuanspace.com/SeeU/

📄 PDF Abstract BibTeX arXiv:2512.03350

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SeeUPO: Sequence-Level Agentic-RL with Convergence Guarantees

2026-02-06 · Tianyi Hu, Qingxu Fu, Yanxi Chen, Zhaoyang Liu 외 arxiv

Reinforcement learning (RL) has emerged as the predominant paradigm for training large language model (LLM)-based AI agents. However, existing backbone RL algorithms lack verified convergence guarantees in agentic scenar…

Reinforcement Learning

Seeing Together: Multi-Robot Cooperative Egocentric Spatial Reasoning with Multimodal Large Language Models

2026-05-18 · Kunyu Peng, Zhikun Zhou, Kailun Yang, Di Wen 외 arxiv

Multimodal Large Language Models (MLLMs) have made substantial progress in egocentric video understanding, but their ability to reason cooperatively from multiple embodied viewpoints remains largely unexplored. We study …

Spatial Reasoning

When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis

2025-01-17 · Ruixuan Zhang, Beichen Wang, Juexiao Zhang, Zilin Bian 외

The increasing availability of traffic videos functioning on a 24/7/365 time scale has the great potential of increasing the spatio-temporal coverage of traffic accidents, which will help improve traffic safety. However,…

Large Language ModelMultimodal Large Language Modelobject-detectionObject Detection+2

From Seeing to Predicting: A Vision-Language Framework for Trajectory Forecasting and Controlled Video Generation

2025-10-01 · Fan Yang, Zhiyang Chen, Yousong Zhu, Xin Li 외 arxiv

Current video generation models produce physically inconsistent motion that violates real-world dynamics. We propose TrajVLM-Gen, a two-stage framework for physics-aware image-to-video generation. First, we employ a Visi…

Trajectory ForecastingTrajectory PredictionVideo Generation

Dynamics-Aligned Latent Imagination in Contextual World Models for Zero-Shot Generalization

2025-08-27 · Frank Röder, Jan Benad, Manfred Eppe, Pradeep Kr. Banerjee arxiv

Real-world reinforcement learning demands adaptation to unseen environmental conditions without costly retraining. Contextual Markov Decision Processes (cMDP) model this challenge, but existing methods often require expl…

Zero-shot GeneralizationReinforcement Learning