paper-with-me

홈 › Papers

One-Shot Learning Meets Depth Diffusion in Multi-Object Videos

2024-08-29 · Anisha Jain

Creating editable videos that depict complex interactions between multiple objects in various artistic styles has long been a challenging task in filmmaking. Progress is often hampered by the scarcity of data sets that contain paired text descriptions and corresponding videos that showcase these interactions. This paper introduces a novel depth-conditioning approach that significantly advances this field by enabling the generation of coherent and diverse videos from just a single text-video pair using a pre-trained depth-aware Text-to-Image (T2I) model. Our method fine-tunes the pre-trained model to capture continuous motion by employing custom-designed spatial and temporal attention mechanisms. During inference, we use the DDIM inversion to provide structural guidance for video generation. This innovative technique allows for continuously controllable depth in videos, facilitating the generation of multiobject interactions while maintaining the concept generation and compositional strengths of the original T2I model across various artistic styles, such as photorealism, animation, and impressionism.

📄 PDF Abstract BibTeX arXiv:2408.16704

Code (0)

등록된 구현이 없습니다.

Tasks

One-Shot LearningVideo Generation

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

Zero-Shot Metric Depth with a Field-of-View Conditioned Diffusion Model

2023-12-20 · Saurabh Saxena, Junhwa Hur, Charles Herrmann, Deqing Sun 외

While methods for monocular depth estimation have made significant strides on standard benchmarks, zero-shot metric depth estimation remains unsolved. Challenges include the joint modeling of indoor and outdoor scenes, w…

DenoisingDepth EstimationMonocular Depth Estimation

OpenDlign: Open-World Point Cloud Understanding with Depth-Aligned Images

2024-04-25 · Ye Mao, Junpeng Jing, Krystian Mikolajczyk

Recent open-world 3D representation learning methods using Vision-Language Models (VLMs) to align 3D point cloud with image-text information have shown superior 3D zero-shot performance. However, CAD-rendered images for …

Representation LearningTransfer LearningZero-shot 3D classificationZero-shot 3D Point Cloud Classification+3

ZeST: Zero-Shot Material Transfer from a Single Image

2024-04-09 · Ta-Ying Cheng, Prafull Sharma, Andrew Markham, Niki Trigoni 외

We propose ZeST, a method for zero-shot material transfer to an object in the input image given a material exemplar image. ZeST leverages existing diffusion adapters to extract implicit material representation from the e…

Appearance TransferObject

DiffuDepGrasp: Diffusion-based Depth Noise Modeling Empowers Sim2Real Robotic Grasping

2025-11-17 · Yingting Zhou, Wenbo Cui, Weiheng Liu, Guixing Chen 외 arxiv

Transferring the depth-based end-to-end policy trained in simulation to physical robots can yield an efficient and robust grasping policy, yet sensor artifacts in real depth maps like voids and noise establish a signific…

Robotic Grasping

Diffusion Meets Few-shot Class Incremental Learning

2025-03-30 · Junsu Kim, Yunhoe Ku, Dongyoon Han, Seungryul Baek

Few-shot class-incremental learning (FSCIL) is challenging due to extremely limited training data; while aiming to reduce catastrophic forgetting and learn new information. We propose Diffusion-FSCIL, a novel approach th…

class-incremental learningClass Incremental LearningFew-Shot Class-Incremental LearningIncremental Learning