paper-with-me

홈 › Papers

Animating Landscape: Self-Supervised Learning of Decoupled Motion and Appearance for Single-Image Video Synthesis

2019-10-16 · Yuki Endo, Yoshihiro Kanamori, Shigeru Kuriyama

Automatic generation of a high-quality video from a single image remains a challenging task despite the recent advances in deep generative models. This paper proposes a method that can create a high-resolution, long-term animation using convolutional neural networks (CNNs) from a single landscape image where we mainly focus on skies and waters. Our key observation is that the motion (e.g., moving clouds) and appearance (e.g., time-varying colors in the sky) in natural scenes have different time scales. We thus learn them separately and predict them with decoupled control while handling future uncertainty in both predictions by introducing latent codes. Unlike previous methods that infer output frames directly, our CNNs predict spatially-smooth intermediate data, i.e., for motion, flow fields for warping, and for appearance, color transfer maps, via self-supervised learning, i.e., without explicitly-provided ground truth. These intermediate data are applied not to each previous output frame, but to the input image only once for each output frame. This design is crucial to alleviate error accumulation in long-term predictions, which is the essential problem in previous recurrent approaches. The output frames can be looped like cinemagraph, and also be controlled directly by specifying latent codes or indirectly via visual annotations. We demonstrate the effectiveness of our method through comparisons with the state-of-the-arts on video prediction as well as appearance manipulation.

📄 PDF Abstract BibTeX arXiv:1910.07192

Code (1)

endo-yuki-t/Animating-Landscape 공식 구현 pytorch

Tasks

Self-Supervised LearningVideo Prediction

Similar Papers 제목 키워드 기반

AHA! Animating Human Avatars in Diverse Scenes with Gaussian Splatting

2025-11-13 · Aymen Mir, Jian Wang, Riza Alp Guler, Chuan Guo 외 arxiv

We present a novel framework for animating humans in 3D scenes using 3D Gaussian Splatting (3DGS), a neural scene representation that has recently achieved state-of-the-art photorealistic results for novel-view synthesis…

Motion SynthesisPoint Clouds

Feature-metric Loss for Self-supervised Learning of Depth and Egomotion

2020-07-21 · ECCV 2020 8 · Chang Shu, Kun Yu, Zhixiang Duan, Kuiyuan Yang

Photometric loss is widely used for self-supervised depth and egomotion estimation. However, the loss landscapes induced by photometric differences are often problematic for optimization, caused by plateau landscapes for…

Depth EstimationMonocular Depth EstimationSelf-Supervised LearningVisual Odometry

Unsupervised Segmentation in Real-World Images via Spelke Object Inference

2022-05-17 · Honglin Chen, Rahul Venkatesh, Yoni Friedman, Jiajun Wu 외

Self-supervised, category-agnostic segmentation of real-world images is a challenging open problem in computer vision. Here, we show how to learn static grouping priors from motion self-supervision by building on the cog…

Image SegmentationOptical Flow EstimationSegmentationSemantic Segmentation

IM-Animation: An Implicit Motion Representation for Identity-decoupled Character Animation

2026-02-07 · Zhufeng Xu, Xuan Gao, Feng-Lin Liu, Haoxian Zhang 외 arxiv

Recent progress in video diffusion models has markedly advanced character animation, which synthesizes motioned videos by animating a static identity image according to a driving video. Explicit methods represent motion …

Weakly-supervised High-fidelity Ultrasound Video Synthesis with Feature Decoupling

2022-07-01 · Jiamin Liang, Xin Yang, Yuhao Huang, Kai Liu 외

Ultrasound (US) is widely used for its advantages of real-time imaging, radiation-free and portability. In clinical practice, analysis and diagnosis often rely on US sequences rather than a single image to obtain dynamic…

Keypoint DetectionVocal Bursts Intensity Prediction