paper-with-me

홈 › Papers

3D Part Guided Image Editing for Fine-Grained Object Understanding

2020-06-01 · CVPR 2020 6 · Zongdai Liu, Feixiang Lu, Peng Wang, Hui Miao, Liangjun Zhang, Ruigang Yang, Bin Zhou

Holistically understanding an object with its 3D movable parts is essential for visual models of a robot to interact with the world. For example, only by understanding many possible part dynamics of other vehicles (e.g., door or trunk opening, taillight blinking for changing lane), a self-driving vehicle can be success in dealing with emergency cases. However, existing visual models tackle rarely on these situations, but focus on bounding box detection. In this paper, we fill this important missing piece in autonomous driving by solving two critical issues. First, for dealing with data scarcity, we propose an effective training data generation process by fitting a 3D car model with dynamic parts to cars in real images. This allows us to directly edit the real images using the aligned 3D parts, yielding effective training data for learning robust deep neural networks (DNNs). Secondly, to benchmark the quality of 3D part understanding, we collected a large dataset in real driving scenario with cars in uncommon states (CUS), i.e. with door or trunk opened etc., which demonstrates that our trained network with edited images largely outperforms other baselines in terms of 2D detection and instance segmentation accuracy.

📄 PDF Abstract BibTeX

Code (1)

zongdai/EditingForDNN 공식 구현 pytorch

Tasks

Autonomous DrivingInstance SegmentationSemantic Segmentation

Similar Papers 제목 키워드 기반

CompBench: Benchmarking Complex Instruction-guided Image Editing

2025-05-18 · Bohan Jia, Wenxuan Huang, Yuntian Tang, Junbo Qiao 외

While real-world applications increasingly demand intricate scene manipulation, existing instruction-guided image editing benchmarks often oversimplify task complexity and lack comprehensive, fine-grained instructions. T…

BenchmarkingInstruction Following

MagGAN: High-Resolution Face Attribute Editing with Mask-Guided Generative Adversarial Network

2020-10-03 · Yi Wei, Zhe Gan, Wenbo Li, Siwei Lyu 외

We present Mask-guided Generative Adversarial Network (MagGAN) for high-resolution face attribute editing, in which semantic facial masks from a pre-trained face parser are used to guide the fine-grained image editing pr…

AttributeGenerative Adversarial NetworkVocal Bursts Intensity Prediction

SpotEdit: Evaluating Visually-Guided Image Editing Methods

2025-08-25 · Sara Ghazanfari, Wei-An Lin, Haitong Tian, Ersin Yumer arxiv

Visually-guided image editing, where edits are conditioned on both visual cues and textual prompts, has emerged as a powerful paradigm for fine-grained, controllable content generation. Although recent generative models …

Image Editing

LocInv: Localization-aware Inversion for Text-Guided Image Editing

2024-05-02 · Chuanming Tang, Kai Wang, Fei Yang, Joost Van de Weijer

Large-scale Text-to-Image (T2I) diffusion models demonstrate significant generation capabilities based on textual prompts. Based on the T2I diffusion models, text-guided image editing research aims to empower users to ma…

Denoisingtext-guided-image-editing

Structure-Preserving Zero-Shot Image Editing via Stage-Wise Latent Injection in Diffusion Models

2025-04-22 · Dasol Jeong, Donggoo Kang, Jiwon Park, Hyebean Lee 외

We propose a diffusion-based framework for zero-shot image editing that unifies text-guided and reference-guided approaches without requiring fine-tuning. Our method leverages diffusion inversion and timestep-specific nu…

Attribute