paper-with-me

Papers

PhysEditBench: A Protocol-Conditioned Benchmark for Dense Physical-Map Prediction with Image Editors

2026-05-13 · Jiaxin Yang, Yu Hou, Muxin Liu, Weixuan Liu, Ze Yuan, Zeming Chen, Zhongrui Wang, Xiaojuan Qi arxiv

Can general-purpose image editors predict physical maps from a single RGB image? General-purpose image editors differ from standard task-specific dense-prediction models: they do not directly take an image and output a physical map. Instead, they must be guided by prompts, examples, or image-based textual cues. To this end, we introduce PhysEditBench, a novel protocol-conditioned benchmark to evaluate and standardize image editors in dense physical-map prediction that covers five targets: depth, normal, albedo, roughness, and metallic maps. For evaluation data, we build a target-dependent benchmark substrate. We use OpenRooms-FF for depth, surface normal, albedo, and roughness, InteriorVerse as an additional source for depth, normal, albedo, and a new procedurally generated source for metallic maps. We curate the data with quality checks, valid-region masks, scene-level sampling, and lighting-based stress subsets to ensure reliable and diverse evaluation. For each target, PhysEditBench defines a fixed protocol that specifies the allowed input, expected output format, and scoring procedure. Each score, therefore, reflects the performance of a model under a specified protocol, rather than its best possible performance under all prompts or interaction modes. Experimental results show that specialized models remain much stronger on depth, normal, and albedo, and stronger image editors can produce more reasonable map-like outputs. For roughness and metallic, image editors can match or outperform specialized baselines on some scalar metrics, but they still suffer from structural errors, sparsity effects, and sensitivity to lighting.

📄 PDF Abstract BibTeX arXiv:2605.13493

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RealX3D: A Physically-Degraded 3D Benchmark for Multi-view Visual Restoration and Reconstruction

2025-12-29 · Shuhong Liu, Chenyu Bao, Ziteng Cui, Yun Liu 외 arxiv

We introduce RealX3D, a real-capture benchmark for multi-view visual restoration and 3D reconstruction under diverse physical degradations. RealX3D groups corruptions into four families, including illumination, scatterin…

3D Reconstruction

Cross-Spectral Dense Correspondence for Multimodal Spectral Medical Imaging

2026-08-28 · Eric L. Wisotzky, Jost Triller, Simon W. Härtl, Oliver T. Bruns 외 arxiv

Precise dense correspondence is a fundamental prerequisite for multimodal spectral imaging systems that fuse disparate wavelength ranges for subsequent analysis in medical and scientific imaging. Corresponding image poin…

Data Augmentation

ACWM-Phys: Investigating Generalized Physical Interaction in Action-Conditioned Video World Models

2026-05-09 · Haotian Xue, Yipu Chen, Liqian Ma, Zelin Zhao 외 arxiv

Action-conditioned world models (ACWMs) have shown strong promise for video prediction and decision-making. However, existing benchmarks are largely restricted to egocentric navigation or narrow, task-specific robotics d…

Video Prediction

Event-Conditioned Diagnostics of Kinematic, Contact, and Object-Permanence Fields in Passive Object-State World Models

2026-06-26 · Yang Liu, Yuming Chen arxiv

World models can predict future physical states, but prediction accuracy alone does not explain how physical information is organized and used inside their latent dynamics. We introduce a controlled diagnostic protocol f…

Geometric Consistency Protocol for Foundation Model Features in Multi-View Satellite Imagery

2026-06-16 · Qiyan Luo, Jie Yang, Yingdong Pi, Lekang Wen 외 arxiv

Standardized evaluation protocols are indispensable for robust benchmarking in remote sensing, particularly as foundation features are increasingly transferred across diverse sensors and complex imaging geometries. In sa…