paper-with-me

Papers

ST-Gen4D: Embedding 4D Spatiotemporal Cognition into World Model for 4D Generation

2026-05-08 · Haonan Wang, Hanyu Zhou, Tao Gu, Luxin Yan arxiv

Generative models have achieved success in producing apparently coherent 2D videos, but remain challenging in the physical world due to lack of 4D spatiotemporal scale. Typically, existing 4D generative models directly embed macro scale constraints to enhance overall spatiotemporal consistency. However, these methods only ensure global appearance coherence and fail to reveal the local dynamics of the physical world. Our insight is that global appearance structure and local dynamic topology empower 4D spatiotemporal cognition, thereby enabling 4D generation with spatiotemporal regularities. In this work, we propose ST-Gen4D, a 4D generation framework with 4D spatiotemporal cognition-based world model. Our model is guided by four key designs: 1) Spatiotemporal representation. We encode various modalities into multiple representations as a feature basis. 2) Spatiotemporal cognition. We sculpture these representations into global appearance graph and local dynamic graph, and fuse them via semantic-bridged spatiotemporal fusion to obtain a 4D cognition graph. 3) Spatiotemporal reasoning. We utilize a world model to derive future state based on the 4D cognition. 4) Spatiotemporal generation. We leverage the derived cognition as condition to guide latent diffusion for 4D Gaussian generation. By deeply integrating 4D intrinsic cognition with generative priors, our model guarantees the structural rationality and topological consistency of 4D generation. Moreover, we propose ST-4D datasets by aggregating public 4D datasets and self-built subset. Extensive experiments demonstrate the superiority of our ST-Gen4D across 3D and 4D generation tasks.

📄 PDF Abstract BibTeX arXiv:2605.07390

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Large-scale Study of Spatiotemporal Representation Learning with a New Benchmark on Action Recognition

2023-03-23 · ICCV 2023 1 · Andong Deng, Taojiannan Yang, Chen Chen

The goal of building a benchmark (suite of datasets) is to provide a unified protocol for fair evaluation and thus facilitate the evolution of a specific area. Nonetheless, we point out that existing protocols of action …

Action RecognitionDomain AdaptationRepresentation LearningSelf-Supervised Learning+2

LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding

2025-05-18 · Hanyu Zhou, Gim Hee Lee

Despite achieving significant progress in 2D image understanding, large multimodal models (LMMs) struggle in the physical world due to the lack of spatial representation. Typically, existing 3D LMMs mainly embed 3D posit…

Scene Understanding

Cross-Modal Learning with 3D Deformable Attention for Action Recognition

2022-12-12 · ICCV 2023 1 · Sangwon Kim, Dasom Ahn, Byoung Chul Ko

An important challenge in vision-based action recognition is the embedding of spatiotemporal features with two or more heterogeneous modalities into a single feature. In this study, we propose a new 3D deformable transfo…

Action Recognition

Towards a geometric understanding of Spatio Temporal Graph Convolution Networks

2023-12-12 · Pratyusha Das, Sarath Shekkizhar, Antonio Ortega

Spatiotemporal graph convolutional networks (STGCNs) have emerged as a desirable model for skeleton-based human action recognition. Despite achieving state-of-the-art performance, there is a limited understanding of the …

Action RecognitionActivity RecognitionDynamic Time WarpingHuman Activity Recognition+1

Dual-Margin Embedding for Fine-Grained Long-Tailed Plant Taxonomy

2025-12-22 · Cheng Yaw Low, Heejoon Koo, Jaewoo Park, Meeyoung Cha arxiv

Taxonomic classification of ecological families, genera, and species underpins biodiversity monitoring and conservation. Existing computer vision methods typically address fine-grained recognition and long-tailed learnin…