paper-with-me

홈 › Papers

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models

2026-04-19 · Hong Jiang, Wensong Song, Zongxin Yang, Ruijie Quan, Yi Yang arxiv

Camera-controllable image editing aims to synthesize novel views of a given scene under varying camera poses while strictly preserving cross-view geometric consistency. However, existing methods typically rely on fragmented geometric guidance, such as only injecting point clouds at the representation level despite models containing multiple levels, and are mainly based on image diffusion models that operate on discrete view mappings. These two limitations jointly lead to geometric drift and structural degradation under continuous camera motion. We observe that while leveraging video models provides continuous viewpoint priors for camera-controllable image editing, they still struggle to form stable geometric understanding if geometric guidance remains fragmented. To systematically address this, we inject unified geometric guidance across three levels that jointly determine the generative output: representation, architecture, and loss function. To this end, we propose UniGeo, a novel camera-controllable editing framework. Specifically, at the representation level, UniGeo incorporates a frame-decoupled geometric reference injection mechanism to provide robust cross-view geometry context. At the architecture level, it introduces geometric anchor attention to align multi-view features. At the loss function level, it proposes a trajectory-endpoint geometric supervision strategy to explicitly reinforce the structural fidelity of target views. Comprehensive experiments across multiple public benchmarks, encompassing both extensive and limited camera motion settings, demonstrate that UniGeo significantly outperforms existing methods in both visual quality and geometric consistency.

📄 PDF Abstract BibTeX arXiv:2604.17565

Code (0)

등록된 구현이 없습니다.

Tasks

Image EditingPoint Clouds

Similar Papers 제목 키워드 기반

UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical Expression

2022-12-06 · Jiaqi Chen, Tong Li, Jinghui Qin, Pan Lu 외

Geometry problem solving is a well-recognized testbed for evaluating the high-level multi-modal reasoning capability of deep models. In most existing works, two main geometry problems: calculation and proving, are usuall…

Geometry Problem SolvingLogical ReasoningMathMathematical Reasoning

UniGeoSeg: Towards Unified Open-World Segmentation for Geospatial Scenes

2025-11-28 · Shuo Ni, Di Wang, He Chen, Haonan Guo 외 arxiv

Instruction-driven segmentation in remote sensing generates masks from guidance, offering great potential for accessible and generalizable applications. However, existing methods suffer from fragmented task formulations …

Zero-shot GeneralizationMulti-Task Learning

UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models

2026-08-05 · Haiyang Zhou, Wangbo Yu, Chaoran Feng, Xunyu Zhou 외 hf

The abundance of casually captured monocular videos and images on social media provides a valuable source for immersive content creation, where generating novel views from such sparse observations can greatly enhance use…

Novel View Synthesis

UniGeo: A Unified 3D Indoor Object Detection Framework Integrating Geometry-Aware Learning and Dynamic Channel Gating

2026-01-30 · Xing Yi, Jinyang Huang, Feng-Qi Cui, Anyang Tong 외 arxiv

The growing adoption of robotics and augmented reality in real-world applications has driven considerable research interest in 3D object detection based on point clouds. While previous methods address unified training ac…

3D Object DetectionPoint Clouds

UniGeo: Taming Video Diffusion for Unified Consistent Geometry Estimation

2025-05-30 · Yang-tian Sun, Xin Yu, Zehuan Huang, Yi-Hua Huang 외

Recently, methods leveraging diffusion model priors to assist monocular geometric estimation (e.g., depth and normal) have gained significant attention due to their strong generalization ability. However, most existing w…

Video Generation