paper-with-me

홈 › Papers

Seeing World Dynamics in a Nutshell

2025-02-05 · Qiuhong Shen, Xuanyu Yi, Mingbao Lin, Hanwang Zhang, Shuicheng Yan, Xinchao Wang

We consider the problem of efficiently representing casually captured monocular videos in a spatially- and temporally-coherent manner. While existing approaches predominantly rely on 2D/2.5D techniques treating videos as collections of spatiotemporal pixels, they struggle with complex motions, occlusions, and geometric consistency due to absence of temporal coherence and explicit 3D structure. Drawing inspiration from monocular video as a projection of the dynamic 3D world, we explore representing videos in their intrinsic 3D form through continuous flows of Gaussian primitives in space-time. In this paper, we propose NutWorld, a novel framework that efficiently transforms monocular videos into dynamic 3D Gaussian representations in a single forward pass. At its core, NutWorld introduces a structured spatial-temporal aligned Gaussian (STAG) representation, enabling optimization-free scene modeling with effective depth and flow regularization. Through comprehensive experiments, we demonstrate that NutWorld achieves high-fidelity video reconstruction quality while enabling various downstream applications in real-time. Demos and code will be available at https://github.com/Nut-World/NutWorld.

📄 PDF Abstract BibTeX arXiv:2502.03465

Code (1)

nut-world/nutworld 공식 구현

Tasks

Video Reconstruction

Similar Papers 제목 키워드 기반

SeeU: Seeing the Unseen World via 4D Dynamics-aware Generation

2025-12-03 · Yu Yuan, Tharindu Wickremasinghe, Zeeshan Nadir, Xijun Wang 외 arxiv

Images and videos are discrete 2D projections of the 4D world (3D space + time). Most visual understanding, prediction, and generation operate directly on 2D observations, leading to suboptimal performance. We propose Se…

Dynamic Review-based Recommenders

2021-10-27 · Kostadin Cvejoski, Ramses J. Sanchez, Christian Bauckhage, Cesar Ojeda

Just as user preferences change with time, item reviews also reflect those same preference changes. In a nutshell, if one is to sequentially incorporate review content knowledge into recommender systems, one is naturally…

Recommendation SystemsReview Generation

NUTSHELL: A Dataset for Abstract Generation from Scientific Talks

2025-02-24 · Maike Züfle, Sara Papi, Beatrice Savoldi, Marco Gaido 외

Scientific communication is receiving increasing attention in natural language processing, especially to help researches access, summarize, and generate content. One emerging application in this area is Speech-to-Abstrac…

Abstract generation

The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World

2023-08-03 · Weiyun Wang, Min Shi, Qingyun Li, Wenhai Wang 외

We present the All-Seeing (AS) project: a large-scale data and model for recognizing and understanding everything in the open world. Using a scalable data engine that incorporates human feedback and efficient models in t…

AllQuestion AnsweringRetrievalText Retrieval

From Seeing to Predicting: A Vision-Language Framework for Trajectory Forecasting and Controlled Video Generation

2025-10-01 · Fan Yang, Zhiyang Chen, Yousong Zhu, Xin Li 외 arxiv

Current video generation models produce physically inconsistent motion that violates real-world dynamics. We propose TrajVLM-Gen, a two-stage framework for physics-aware image-to-video generation. First, we employ a Visi…

Trajectory ForecastingTrajectory PredictionVideo Generation