paper-with-me

홈 › Papers

SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing

2026-04-06 · Yicheng Xiao, Wenhu Zhang, Lin Song, Yukang Chen, Wenbo Li, Nan Jiang, Tianhe Ren, Haokun Lin, Wei Huang, Haoyang Huang, Xiu Li, Nan Duan, Xiaojuan Qi arxiv

Image spatial editing performs geometry-driven transformations, allowing precise control over object layout and camera viewpoints. Current models are insufficient for fine-grained spatial manipulations, motivating a dedicated assessment suite. Our contributions are listed: (i) We introduce SpatialEdit-Bench, a complete benchmark that evaluates spatial editing by jointly measuring perceptual plausibility and geometric fidelity via viewpoint reconstruction and framing analysis. (ii) To address the data bottleneck for scalable training, we construct SpatialEdit-500k, a synthetic dataset generated with a controllable Blender pipeline that renders objects across diverse backgrounds and systematic camera trajectories, providing precise ground-truth transformations for both object- and camera-centric operations. (iii) Building on this data, we develop SpatialEdit-16B, a baseline model for fine-grained spatial editing. Our method achieves competitive performance on general editing while substantially outperforming prior methods on spatial manipulation tasks. All resources will be made public at https://github.com/EasonXiao-888/SpatialEdit.

📄 PDF Abstract BibTeX arXiv:2604.04911

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RoadBench: Benchmarking MLLMs on Fine-Grained Spatial Understanding and Reasoning under Urban Road Scenarios

2025-11-22 · Jun Zhang, Xin Zhang, Jie Feng, Long Chen 외 arxiv

Multimodal large language models (MLLMs) have demonstrated powerful capabilities in general spatial understanding and reasoning. However, their fine-grained spatial understanding and reasoning capabilities in complex urb…

BOP-ASK: Object-Interaction Reasoning for Vision-Language Models

2025-11-20 · Vineet Bhat, Sungsu Kim, Valts Blukis, Greg Heinrich 외 arxiv

Vision Language Models (VLMs) have achieved impressive performance on spatial reasoning benchmarks, yet these evaluations mask critical weaknesses in understanding object interactions. Current benchmarks test high level …

Trajectory PlanningSpatial ReasoningPose Estimation

Foundation Models for Remote Sensing: An Analysis of MLLMs for Object Localization

2025-04-14 · Darryl Hannan, John Cooper, Dylan White, Timothy Doster 외

Multimodal large language models (MLLMs) have altered the landscape of computer vision, obtaining impressive results across a wide range of tasks, especially in zero-shot settings. Unfortunately, their strong performance…

BenchmarkingEarth ObservationImage CaptioningObject Localization+2

DataEvolver: Let Your Data Build and Improve Itself via Goal-Driven Loop Agents

2026-05-03 · Qisong Zhang, Wenzhuo Wu, Zhuangzhuang Jia, Yunhao Yang 외 arxiv

Constructing controllable visual data is a major bottleneck for image editing and multimodal understanding. Useful supervision is rarely produced by a single rendering pass; instead it emerges through iterative generatio…

Image Editing

Beyond Binary Success: A Diagnostic Meta-Evaluation Framework for Fine-Grained Manipulation

2026-05-19 · He-Yang Xu, Pengyuan Zhang, Zongyuan Ge, Xiaoshuai Hao 외 arxiv

Fine-grained manipulation marks a regime where global scene context no longer suffices, and success hinges on the tight coupling of local attribute grounding, high-fidelity spatial perception, and constraint-respecting m…