paper-with-me

홈 › Papers

Fine-Grained Human Pose Editing Assessment via Layer-Selective MLLMs

2026-01-15 · Ningyu Sun, Zhaolin Cai, Zitong Xu, Peihang Chen, Huiyu Duan, Yichao Yan, Xiongkuo Min, Xiaokang Yang arxiv

Text-guided human pose editing has gained significant traction in AIGC applications. However,it remains plagued by structural anomalies and generative artifacts. Existing evaluation metrics often isolate authenticity detection from quality assessment, failing to provide fine-grained insights into pose-specific inconsistencies. To address these limitations, we introduce HPE-Bench, a specialized benchmark comprising 1,700 standardized samples from 17 state-of-the-art editing models, offering both authenticity labels and multi-dimensional quality scores. Furthermore, we propose a unified framework based on layer-selective multimodal large language models (MLLMs). By employing contrastive LoRA tuning and a novel layer sensitivity analysis (LSA) mechanism, we identify the optimal feature layer for pose evaluation. Our framework achieves superior performance in both authenticity detection and multi-dimensional quality regression, effectively bridging the gap between forensic detection and quality assessment.

📄 PDF Abstract BibTeX arXiv:2601.10369

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Human-Aligned MLLM Judges for Fine-Grained Image Editing Evaluation: A Benchmark, Framework, and Analysis

2026-02-13 · Runzhou Liu, Hailey Weingord, Sejal Mittal, Prakhar Dungarwal 외 arxiv

Evaluating image editing models remains challenging due to the coarse granularity and limited interpretability of traditional metrics, which often fail to capture aspects important to human perception and intent. Such me…

Image Editing

Fine-grained Hallucination Detection and Editing for Language Models

2024-01-12 · Abhika Mishra, Akari Asai, Vidhisha Balachandran, Yizhong Wang 외

Large language models (LMs) are prone to generate factual errors, which are often called hallucinations. In this paper, we introduce a comprehensive taxonomy of hallucinations and argue that hallucinations manifest in di…

HallucinationRetrieval

Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling

2026-05-13 · Xuehai Bai, Yang Shi, Yi-Fan Zhang, Xuanyu Zhu 외 arxiv

Recent image editing models have achieved remarkable progress in instruction following, multimodal understanding, and complex visual editing. However, existing benchmarks often fail to faithfully reflect human judgment, …

Instruction FollowingVisual ReasoningImage Editing

SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing

2026-04-06 · Yicheng Xiao, Wenhu Zhang, Lin Song, Yukang Chen 외 arxiv

Image spatial editing performs geometry-driven transformations, allowing precise control over object layout and camera viewpoints. Current models are insufficient for fine-grained spatial manipulations, motivating a dedi…

FED-Bench: A Cross-Granular Benchmark for Disentangled Evaluation of Facial Expression Editing

2026-03-31 · Fengjian Xue, Xuecheng Wu, Heli Sun, Yunyun Shi 외 arxiv

Facial expression image editing requires fine-grained control to strictly preserve human identity and background while precisely manipulating expression. However, existing editing benchmarks primarily focus on general sc…

Instruction FollowingImage Editing