paper-with-me

홈 › Papers

MOF-Sleuth: Tool-Grounded Reward Alignment for Explainable Fine-Grained MOF CIF Auditing

2026-07-22 · Yu Liu, Zhiwei Yang, Diandian Guo, Kun Peng, Fangfang Yuan, Cong Cao, Chaozhuo Li, Zhiyuan Ma, Yanbing Liu, Guobin Zhao arxiv

Large metal-organic framework (MOF) databases support simulation, screening, and machine learning through crystallographic information files (CIFs). Subtle chemical and structural errors in these inputs can compromise downstream results and hinder manual inspection. LLM advances in computational chemistry offer paths beyond predictive screening toward fine-grained diagnosis with evidence-grounded explanations. However, two challenges remain: (i) limited fine-grained attribution: MOF-specific validators and machine-learning models scale detection but provide fixed checks, readiness scores, or coarse labels rather than evidence-grounded explanations; and (ii) unreliable CIF reasoning: direct LLM auditing is costly and unreliable because chemical evidence is implicit across atom-site records and requires geometric, connectivity, occupancy, and charge calculations. Both stem from weak coupling between chemical evidence and language-model explanation. We introduce MOF-Sleuth, a reinforcement-guided CIF auditing agent with two modules: a deterministic Forensic Lab and a Sleuth reasoning engine. The Lab derives composition, geometry, connectivity, occupancy, coordination, and charge evidence, and Sleuth uses this evidence to produce an evidence-grounded explanation, error types, and a binary decision. Reward-guided reinforcement learning (RL) turns tool measurements into chemical explanation-level supervision, rewarding not only the final answer but also cited chemical evidence and evidence-supported diagnoses. We introduce Chemically Grounded Diagnosis (Chem-GD), a metric that assesses whether a correct diagnosis is explained by factual, relevant CIF-derived evidence. Across four benchmarks, MOF-Sleuth establishes state-of-the-art performance among LLM-based approaches and MOF-specific machine-learning methods, demonstrating gains in detection, attribution, and grounded explanation quality.

📄 PDF Abstract BibTeX arXiv:2607.19935

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

EditSleuth: A Dataset of Grounded Reasoning Chains for Image-Edit Forensics

2026-05-09 · Van-Loc Nguyen, AprilPyone MaungMaung, Minh-Triet Tran, Isao Echizen arxiv

Forensic analysis of AI-edited images requires more than binary real-versus-fake prediction: a useful system should localize the edit, identify its semantic type, and ground its decisions in visual evidence. Existing ima…

Image Manipulation

AcrosticSleuth: Probabilistic Identification and Ranking of Acrostics in Multilingual Corpora

2024-08-08 · Aleksandr Fedchin, Isabel Cooperman, Pramit Chaudhuri, Joseph P. Dexter

For centuries, writers have hidden messages in their texts as acrostics, where initial letters of consecutive lines or paragraphs form meaningful words or phrases. Scholars searching for acrostics manually can only focus…

Binary Classification

TextSleuth: A New Dataset and Baseline for Scene Text Manipulation Detection

2024-08-07 · Conference on Multimedia Information Processing and Retrieval 2024 8 · Abhineet Kumar Pandey, Ming-Ching Chang Xin, Li

With the rise of digital content on social media and the advancement of image editing tools, tampering with scene text has become a serious concern. Scene text manipulation detection (STMD) is a kind of image manipulatio…

Image ManipulationImage Manipulation Detection

Track, Rank, Crack: Epistemic Working Memory Scales Multi-Hop Reasoning in Language Agents

2026-07-14 · Ning Liu arxiv

Language agents that interleave reasoning and tool use degrade sharply as reasoning chains lengthen, even when each individual step is easy. We trace this to context dilution: an agent's investigative state (what it has …

Which Models Are Our Models Built On? Auditing Invisible Dependencies in Modern LLMs

2026-06-10 · Sanjay Adhikesaven, Haoxiang Sun, Sewon Min arxiv

Modern LLM training pipelines increasingly rely on other models to generate data, filter corpora, judge outputs, and guide development decisions. These dependencies are recursive: a model may depend on an upstream artifa…

Information Extraction