paper-with-me

홈 › Papers

Learning Semantic-Aware Dynamics for Video Prediction

2021-04-20 · CVPR 2021 1 · Xinzhu Bei, Yanchao Yang, Stefano Soatto

We propose an architecture and training scheme to predict video frames by explicitly modeling dis-occlusions and capturing the evolution of semantically consistent regions in the video. The scene layout (semantic map) and motion (optical flow) are decomposed into layers, which are predicted and fused with their context to generate future layouts and motions. The appearance of the scene is warped from past frames using the predicted motion in co-visible regions; dis-occluded regions are synthesized with content-aware inpainting utilizing the predicted scene layout. The result is a predictive model that explicitly represents objects and learns their class-specific motion, which we evaluate on video prediction benchmarks.

📄 PDF Abstract BibTeX arXiv:2104.09762

Code (0)

등록된 구현이 없습니다.

Tasks

Optical Flow EstimationPredictionVideo Prediction

Methods 이 논문이 사용한 방법론

Inpainting Train a convolutional neural network to generate the contents of an arbitrary image region conditioned on its surroundings.

Similar Papers 제목 키워드 기반

DynaRend: Learning 3D Dynamics via Masked Future Rendering for Robotic Manipulation

2025-10-28 · Jingyi Tian, Le Wang, Sanping Zhou, Sen Wang 외 arxiv

Learning generalizable robotic manipulation policies remains a key challenge due to the scarcity of diverse real-world training data. While recent approaches have attempted to mitigate this through self-supervised repres…

Representation LearningVideo Prediction

DynaMind: Reconstructing Dynamic Visual Scenes from EEG by Aligning Temporal Dynamics and Multimodal Semantics to Guided Diffusion

2025-09-01 · Junxiang Liu, Junming Lin, Jiangtong Li, Jie Li arxiv

Reconstruction dynamic visual scenes from electroencephalography (EEG) signals remains a primary challenge in brain decoding, limited by the low spatial resolution of EEG, a temporal mismatch between neural recordings an…

Video ReconstructionBrain Decoding

Flowing from Reasoning to Motion: Learning 3D Hand Trajectory Prediction from Egocentric Human Interaction Videos

2025-12-18 · Mingfei Chen, Yifan Wang, Zhengqin Li, Homanga Bharadhwaj 외 arxiv

Prior works on 3D hand trajectory prediction are constrained by datasets that decouple motion from semantic supervision and by models that weakly link reasoning and action. To address these, we first present the EgoMAN d…

Trajectory Prediction

Unbiased Video Scene Graph Generation via Visual and Semantic Dual Debiasing

2025-03-01 · CVPR 2025 1 · Yanjun Li, Zhaoyang Li, Honghui Chen, Lizhi Xu

Video Scene Graph Generation (VidSGG) aims to capture dynamic relationships among entities by sequentially analyzing video frames and integrating visual and semantic information. However, VidSGG is challenged by signific…

Graph GenerationScene Graph GenerationTripletVideo scene graph generation

Semantic-Aware, Physics-Informed, Geometry-Grounded Weather Video Synthesis

2026-06-27 · Chenghao Qian, Nedko Savov, Lingdong Kong, Yeying Jin 외 arxiv

Weather synthesis aims to add weather effects to input videos while preserving scene identity, structure, and motion. The key limitation of existing methods is the lack of diversity in weather appearance and effective co…

Semantic SegmentationAutonomous Driving