paper-with-me

홈 › Papers

Locate-Then-Examine: Grounded Region Reasoning Improves Detection of AI-Generated Images

2025-10-05 · Yikun Ji, Yan Hong, Bowen Deng, Jun Lan, Huijia Zhu, Weiqiang Wang, Liqing Zhang, Jianfu Zhang arxiv

The rapid growth of AI-generated imagery has blurred the boundary between real and synthetic content, raising practical concerns for digital integrity. Vision-language models (VLMs) can provide natural language explanations, but standard one-pass classifiers often miss subtle artifacts in high-quality synthetic images and offer limited grounding in the pixels. We propose Locate-Then-Examine (LTE), a two-stage VLM-based forensic framework that first localizes suspicious regions and then re-examines these crops together with the full image to refine the real vs. AI-generated verdict and its explanation. LTE explicitly links each decision to localized visual evidence through region proposals and region-aware reasoning. To support training and evaluation, we introduce TRACE, a dataset of 20,000 real and high-quality synthetic images with region-level annotations and automatically generated forensic explanations, constructed by a VLM-based pipeline with additional consistency checks and quality control. Across TRACE and multiple external benchmarks, LTE achieves competitive accuracy and improved robustness while providing human-understandable, region-grounded explanations suitable for forensic deployment.

📄 PDF Abstract BibTeX arXiv:2510.04225

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

VisionCoach: Reinforcing Grounded Video Reasoning via Visual-Perception Prompting

2026-03-15 · Daeun Lee, Shoubin Yu, Yue Zhang, Mohit Bansal arxiv

Video reasoning requires models to locate and track question-relevant evidence across frames. While reinforcement learning (RL) with verifiable rewards improves accuracy, it still struggles to achieve reliable spatio-tem…

Reinforcement Learning

Localize, Then Reason: Visual Latent Structural Reasoning for Molecular Properties and Edits

2026-08-13 · Xingqiao Lin, Junmei Wang, Haocheng Tang arxiv

Local chemical perception and property reasoning are both essential for understanding how molecular structure determines properties. Current LLM-based chemical reasoning methods either receive SMILES/molecular images tog…

DocCogito: Aligning Layout Cognition and Step-Level Grounded Reasoning for Document Understanding

2026-03-08 · Yuchuan Wu, Minghan Zhuo, Teng Fu, Mengyang Zhao 외 arxiv

Document understanding with multimodal large language models (MLLMs) requires not only accurate answers but also explicit, evidence-grounded reasoning, especially in high-stakes scenarios. However, current document MLLMs…

Towards Understanding Visual Grounding in Visual Language Models

2025-09-12 · Georgios Pantazopoulos, Eda B. Özyiğit arxiv

Visual grounding refers to the ability of a model to identify a region within some visual input that matches a textual description. Consequently, a model equipped with visual grounding capabilities can target a wide rang…

multimodal generationReferring ExpressionVisual Grounding

MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding

2026-04-10 · Henry Zheng, Chenyue Fang, Rui Huang, Siyuan Wei 외 arxiv

Vision-language models (VLMs) have achieved strong performance in multimodal understanding and reasoning, yet grounded reasoning in 3D scenes remains underexplored. Effective 3D reasoning hinges on accurate grounding: to…

Zero-shot Generalization