paper-with-me

홈 › Papers

GenVideo: One-shot Target-image and Shape Aware Video Editing using T2I Diffusion Models

2024-04-18 · Sai Sree Harsha, Ambareesh Revanur, Dhwanit Agarwal, Shradha Agrawal

Video editing methods based on diffusion models that rely solely on a text prompt for the edit are hindered by the limited expressive power of text prompts. Thus, incorporating a reference target image as a visual guide becomes desirable for precise control over edit. Also, most existing methods struggle to accurately edit a video when the shape and size of the object in the target image differ from the source object. To address these challenges, we propose "GenVideo" for editing videos leveraging target-image aware T2I models. Our approach handles edits with target objects of varying shapes and sizes while maintaining the temporal consistency of the edit using our novel target and shape aware InvEdit masks. Further, we propose a novel target-image aware latent noise correction strategy during inference to improve the temporal consistency of the edits. Experimental analyses indicate that GenVideo can effectively handle edits with objects of varying shapes, where existing approaches fail.

📄 PDF Abstract BibTeX arXiv:2404.12541

Code (0)

등록된 구현이 없습니다.

Tasks

Video Editing

Methods 이 논문이 사용한 방법론

AWARE We propose to theoretically and empirically examine the effect of incorporating weighting schemes into walk-aggregating GNNs. To this end, we propose a simple, interpretable, and…
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Few-shot 3D Shape Generation

2023-05-19 · Jingyuan Zhu, Huimin Ma, Jiansheng Chen, Jian Yuan

Realistic and diverse 3D shape generation is helpful for a wide variety of applications such as virtual reality, gaming, and animation. Modern generative models, such as GANs and diffusion models, learn from large-scale …

3D Shape GenerationDiversityDomain AdaptationImage Generation

ShapeWords: Guiding Text-to-Image Synthesis with 3D Shape-Aware Prompts

2024-12-03 · CVPR 2025 1 · Dmitry Petrov, Pradyumn Goyal, Divyansh Shivashok, Yuanming Tao 외

We introduce ShapeWords, an approach for synthesizing images based on 3D shape guidance and text prompts. ShapeWords incorporates target 3D shape information within specialized tokens embedded together with the input tex…

Image Generation

Delving into Shape-aware Zero-shot Semantic Segmentation

2023-04-17 · CVPR 2023 1 · Xinyu Liu, Beiwen Tian, Zhen Wang, Rui Wang 외

Thanks to the impressive progress of large-scale vision-language pretraining, recent recognition models can classify arbitrary objects in a zero-shot and open-set manner, with a surprisingly high accuracy. However, trans…

Image SegmentationSegmentationSemantic SegmentationZero-Shot Semantic Segmentation

DeMamba: AI-Generated Video Detection on Million-Scale GenVideo Benchmark

2024-05-30 · Haoxing Chen, Yan Hong, Zizheng Huang, Zhuoer Xu 외

Recently, video generation techniques have advanced rapidly. Given the popularity of video content on social media platforms, these models intensify concerns about the spread of fake information. Therefore, there is a gr…

DeepFake DetectionMambaVideo ClassificationVideo Generation+2

Few-shot Shape Recognition by Learning Deep Shape-aware Features

2023-12-03 · Wenlong Shi, Changsheng Lu, Ming Shao, Yinjie Zhang 외

Traditional shape descriptors have been gradually replaced by convolutional neural networks due to their superior performance in feature extraction and classification. The state-of-the-art methods recognize object shapes…

Image Reconstruction