paper-with-me

홈 › Papers

Render In-between: Motion Guided Video Synthesis for Action Interpolation

2021-11-01 · Hsuan-I Ho, Xu Chen, Jie Song, Otmar Hilliges

Upsampling videos of human activity is an interesting yet challenging task with many potential applications ranging from gaming to entertainment and sports broadcasting. The main difficulty in synthesizing video frames in this setting stems from the highly complex and non-linear nature of human motion and the complex appearance and texture of the body. We propose to address these issues in a motion-guided frame-upsampling framework that is capable of producing realistic human motion and appearance. A novel motion model is trained to inference the non-linear skeletal motion between frames by leveraging a large-scale motion-capture dataset (AMASS). The high-frame-rate pose predictions are then used by a neural rendering pipeline to produce the full-frame output, taking the pose and background consistency into consideration. Our pipeline only requires low-frame-rate videos and unpaired human motion data but does not require high-frame-rate videos for training. Furthermore, we contribute the first evaluation dataset that consists of high-quality and high-frame-rate videos of human activities for this task. Compared with state-of-the-art video interpolation techniques, our method produces in-between frames with better quality and accuracy, which is evident by state-of-the-art results on pixel-level, distributional metrics and comparative user evaluations. Our code and the collected dataset are available at https://git.io/Render-In-Between.

📄 PDF Abstract BibTeX arXiv:2111.01029

Code (1)

azuxmioy/Render-In-Between 공식 구현 pytorch

Tasks

Neural Rendering

Similar Papers 제목 키워드 기반

Stereo-Talker: Audio-driven 3D Human Synthesis with Prior-Guided Mixture-of-Experts

2024-10-31 · Xiang Deng, Youxin Pang, Xiaochen Zhao, Chao Xu 외

This paper introduces Stereo-Talker, a novel one-shot audio-driven human video synthesis system that generates 3D talking videos with precise lip synchronization, expressive body gestures, temporally consistent photo-rea…

Language ModelingLanguage ModellingLarge Language ModelMixture-of-Experts+1

DynaVid: Learning to Generate Highly Dynamic Videos using Synthetic Motion Data

2026-04-02 · Wonjoon Jin, Jiyun Won, Janghyeok Han, Qi Dai 외 arxiv

Despite recent progress, video diffusion models still struggle to synthesize realistic videos involving highly dynamic motions or requiring fine-grained motion controllability. A central limitation lies in the scarcity o…

Lang2Motion: Bridging Language and Motion through Joint Embedding Spaces

2025-12-11 · Bishoy Galoaa, Xiangyu Bai, Sarah Ostadabbas arxiv

We present Lang2Motion, a framework for language-guided point trajectory generation by aligning motion manifolds with joint embedding spaces. Unlike prior work focusing on human motion or video synthesis, we generate exp…

Action RecognitionVideo GenerationPoint TrackingStyle Transfer

A Good Image Generator Is What You Need for High-Resolution Video Synthesis

2021-04-30 · ICLR 2021 1 · Yu Tian, Jian Ren, Menglei Chai, Kyle Olszewski 외

Image and video synthesis are closely related areas aiming at generating content from noise. While rapid progress has been demonstrated in improving image-based models to handle large resolutions, high-quality renderings…

Video Generation

CPSL: Representing Volumetric Video via Content-Promoted Scene Layers

2025-11-18 · Kaiyuan Hu, Yili Jin, Junhua Liu, Xize Duan 외 arxiv

Volumetric video enables immersive and interactive visual experiences by supporting free viewpoint exploration and realistic motion parallax. However, existing volumetric representations from explicit point clouds to imp…

3D ReconstructionPoint Clouds