paper-with-me

Papers

Unified Editing of Panorama, 3D Scenes, and Videos Through Disentangled Self-Attention Injection

2024-05-27 · Gihyun Kwon, Jangho Park, Jong Chul Ye

While text-to-image models have achieved impressive capabilities in image generation and editing, their application across various modalities often necessitates training separate models. Inspired by existing method of single image editing with self attention injection and video editing with shared attention, we propose a novel unified editing framework that combines the strengths of both approaches by utilizing only a basic 2D image text-to-image (T2I) diffusion model. Specifically, we design a sampling method that facilitates editing consecutive images while maintaining semantic consistency utilizing shared self-attention features during both reference and consecutive image sampling processes. Experimental results confirm that our method enables editing across diverse modalities including 3D scenes, videos, and panorama images.

📄 PDF Abstract BibTeX arXiv:2405.16823

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationVideo Editing

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

VidPanos: Generative Panoramic Videos from Casual Panning Videos

2024-10-17 · Jingwei Ma, Erika Lu, Roni Paiss, Shiran Zada 외

Panoramic image stitching provides a unified, wide-angle view of a scene that extends beyond the camera's field of view. Stitching frames of a panning video into a panoramic photograph is a well-understood problem for st…

Image StitchingVideo Generation

StyleLight: HDR Panorama Generation for Lighting Estimation and Editing

2022-07-29 · Guangcong Wang, Yinuo Yang, Chen Change Loy, Ziwei Liu

We present a new lighting estimation and editing framework to generate high-dynamic-range (HDR) indoor panorama lighting from a single limited field-of-view (LFOV) image captured by low-dynamic-range (LDR) cameras. Exist…

Lighting Estimation

Collaborative Score Distillation for Consistent Visual Editing

2023-09-21 · NeurIPS 2023 11

Generative priors of large-scale text-to-image diffusion models enable a wide range of new generation and editing applications on diverse visual modalities. However, when adapting these priors to complex visual modalitie…

More than the Sum: Panorama-Language Models for Adverse Omni-Scenes

2026-03-10 · Weijia Fan, Ruiping Liu, Jiale Wei, Yufan Chen 외 arxiv

Existing vision-language models (VLMs) are tailored for pinhole imagery, stitching multiple narrow field-of-view inputs to piece together a complete omni-scene understanding. Yet, such multi-view perception overlooks the…

Scene Understanding

OmniX: From Unified Panoramic Generation and Perception to Graphics-Ready 3D Scenes

2025-10-30 · Yukun Huang, Jiwen Yu, Yanning Zhou, Jianan Wang 외 arxiv

There are two prevalent ways to constructing 3D scenes: procedural generation and 2D lifting. Among them, panorama-based 2D lifting has emerged as a promising technique, leveraging powerful 2D generative priors to produc…

Scene Generation