paper-with-me

홈 › Papers

Neural USD: An object-centric framework for iterative editing and control

2025-10-28 · Alejandro Escontrela, Shrinu Kushagra, Sjoerd van Steenkiste, Yulia Rubanova, Aleksander Holynski, Kelsey Allen, Kevin Murphy, Thomas Kipf arxiv

Amazing progress has been made in controllable generative modeling, especially over the last few years. However, some challenges remain. One of them is precise and iterative object editing. In many of the current methods, trying to edit the generated image (for example, changing the color of a particular object in the scene or changing the background while keeping other elements unchanged) by changing the conditioning signals often leads to unintended global changes in the scene. In this work, we take the first steps to address the above challenges. Taking inspiration from the Universal Scene Descriptor (USD) standard developed in the computer graphics community, we introduce the "Neural Universal Scene Descriptor" or Neural USD. In this framework, we represent scenes and objects in a structured, hierarchical manner. This accommodates diverse signals, minimizes model-specific constraints, and enables per-object control over appearance, geometry, and pose. We further apply a fine-tuning approach which ensures that the above control signals are disentangled from one another. We evaluate several design considerations for our framework, demonstrating how Neural USD enables iterative and incremental workflows. More information at: https://escontrela.me/neural_usd .

📄 PDF Abstract BibTeX arXiv:2510.23956

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FlexEdit: Flexible and Controllable Diffusion-based Object-centric Image Editing

2024-03-27 · Trong-Tung Nguyen, Duc-Anh Nguyen, Anh Tran, Cuong Pham

Our work addresses limitations seen in previous approaches for object-centric editing problems, such as unrealistic results due to shape discrepancies and limited control in object replacement or insertion. To this end, …

DenoisingObjecttext-guided-image-editing

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation

2026-04-13 · Yuqian Yuan, Wenqiao Zhang, Juekai Lin, Yu Zhong 외 arxiv

Large Multimodal Models (LMMs) have achieved remarkable progress in general-purpose vision--language understanding, yet they remain limited in tasks requiring precise object-level grounding, fine-grained spatial reasonin…

Scene UnderstandingSpatial Reasoning

VASE: Object-Centric Appearance and Shape Manipulation of Real Videos

2024-01-04 · Elia Peruzzo, Vidit Goel, Dejia Xu, Xingqian Xu 외

Recently, several works tackled the video editing task fostered by the success of large-scale text-to-image generative models. However, most of these methods holistically edit the frame using the text, exploiting the pri…

Video Editing

Learning Object-Centric Representations Based on Slots in Real World Scenarios

2025-09-29 · Adil Kaan Akan arxiv

A central goal in AI is to represent scenes as compositions of discrete objects, enabling fine-grained, controllable image and video generation. Yet leading diffusion models treat images holistically and rely on text con…

Unsupervised Video Object SegmentationVideo GenerationImage Generation

Scene Editing as Teleoperation: A Case Study in 6DoF Kit Assembly

2021-10-09 · Yulong Li, Shubham Agrawal, Jen-Shuo Liu, Steven K. Feiner 외

Studies in robot teleoperation have been centered around action specifications -- from continuous joint control to discrete end-effector pose control. However, these robot-centric interfaces often require skilled operato…