paper-with-me

홈 › Papers

Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations

2026-09-25 · Yonghong Zhang, Yong Xie, Isabel M. Parra, Ricardo Correia hf

Difference-in-differences (DID) studies are widely used to evaluate climate policy, but assessing the evidence supporting their identification assumptions remains challenging. We introduce ARGUS, a structured language-model pipeline that audits reported evidence against an eleven-dimension assumption-implication-evidence rubric and abstains when relevant evidence cannot be retrieved. We evaluate ARGUS using injected flaws, economics papers, and a small pilot with reconciled labels. On the 11-flaw benchmark, ARGUS detects 73% of planted flaws, compared with 18% for a keyword-based pipeline. Across 26 economics papers, ARGUS abstains on about 40% of paper-dimension assessments for lack of retrievable evidence. In a five-paper pilot with labels reconciled by two annotators, it assigns a higher risk level than the labels on 25 of the 33 assessments it completes. A rule fixed before the labels arrived removes most of this in-sample; weighted agreement stays low. ARGUS provides evidence-linked risk reports that localize potential weaknesses for expert review, without adjudicating causal claims. Code and data: https://github.com/yonghongzhang-io/ARGUS

📄 PDF Abstract BibTeX arXiv:2609.30867

Code (2)

Tavish9/awesome-daily-AI-arxiv ★ 121
Valiant-Cat/hfpaper

Similar Papers 제목 키워드 기반

Martingale Doppelgänger-Eval: An Identification Framework for Auditing Candlestick Understanding in Vision-Language Models

2026-06-16 · Ziyao Wang arxiv

We introduce Martingale Doppelgänger-Eval, a public shadow-market benchmark for auditing whether vision-language models (VLMs) use candlestick evidence rather than extrapolate past trends. The central difficulty is ident…

MOF-Sleuth: Tool-Grounded Reward Alignment for Explainable Fine-Grained MOF CIF Auditing

2026-07-22 · Yu Liu, Zhiwei Yang, Diandian Guo, Kun Peng 외 arxiv

Large metal-organic framework (MOF) databases support simulation, screening, and machine learning through crystallographic information files (CIFs). Subtle chemical and structural errors in these inputs can compromise do…

Reinforcement Learning

GCA Framework: A GCC Countries-Grounded Dataset and Agentic Pipeline for Climate Decision Support

2026-04-14 · Muhammad Umer Sheikh, Khawar Shehzad, Salman Khan, Fahad Shahbaz Khan 외 arxiv

Climate decision-making in the GCC states increasingly demands systems that can translate heterogeneous scientific and policy evidence into actionable guidance, yet general-purpose large language models (LLMs) remain wea…

EvidenceLens: A Claim-Evidence Matrix for Auditing Financial Question Answering

2026-06-19 · Fengchen Gu, Xiaotian Ren, Zhengyong Jiang, Zhilu Zhang 외 arxiv

Large language models are increasingly used to answer questions over annual reports, earnings decks, and analyst notes, yet their outputs remain difficult to verify in high-stakes financial workflows. A fluent answer can…

Question Answering

Adults in the room? The auditor and dividends in small firms: Evidence from a natural experiment

2023-01-26 · Hakim Lyngstadås, Johannes Mauritzen

We examine the effect of auditing on dividends in small private firms. We hypothesize that auditing can constrain dividends by way of promoting accounting conservatism. We use register data on private Norwegian firms and…

regression