paper-with-me

Papers

Project Ariadne: A Structural Causal Framework for Auditing Faithfulness in LLM Agents

2026-01-05 · Sourena Khanzadeh arxiv

As Large Language Model (LLM) agents are increasingly tasked with high-stakes autonomous decision-making, the transparency of their reasoning processes has become a critical safety concern. While \textit{Chain-of-Thought} (CoT) prompting allows agents to generate human-readable reasoning traces, it remains unclear whether these traces are \textbf{faithful} generative drivers of the model's output or merely \textbf{post-hoc rationalizations}. We introduce \textbf{Project Ariadne}, a novel XAI framework that utilizes Structural Causal Models (SCMs) and counterfactual logic to audit the causal integrity of agentic reasoning. Unlike existing interpretability methods that rely on surface-level textual similarity, Project Ariadne performs \textbf{hard interventions} ($do$-calculus) on intermediate reasoning nodes -- systematically inverting logic, negating premises, and reversing factual claims -- to measure the \textbf{Causal Sensitivity} ($φ$) of the terminal answer. Our empirical evaluation of state-of-the-art models reveals a persistent \textit{Faithfulness Gap}. We define and detect a widespread failure mode termed \textbf{Causal Decoupling}, where agents exhibit a violation density ($ρ$) of up to $0.77$ in factual and scientific domains. In these instances, agents arrive at identical conclusions despite contradictory internal logic, proving that their reasoning traces function as "Reasoning Theater" while decision-making is governed by latent parametric priors. Our findings suggest that current agentic architectures are inherently prone to unfaithful explanation, and we propose the Ariadne Score as a new benchmark for aligning stated logic with model action.

📄 PDF Abstract BibTeX arXiv:2601.02314

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Project Ariadne: Prompt-Conditioned Route Generation for Synthesis Planning

2026-06-23 · Anton Morgunov, Victor S. Batista arxiv

Retrosynthetic planning seeks to connect a target molecule to commercially available starting materials through a multistep route. Classical planners construct such routes by iteratively applying single-step reaction mod…

ISAAC: Auditing Causal Reasoning in Deep Models for Drug-Target Interaction

2026-05-03 · Barbara Tarantino, Sun Kim, Yijingxiu Lu, Paolo Giudici arxiv

Deep learning models for drug--target interaction (DTI) prediction often achieve strong benchmark performance without necessarily relying on mechanistically meaningful molecular features, a limitation that standard accur…

ARIADNE: A Perception-Reasoning Synergy Framework for Trustworthy Coronary Angiography Analysis

2026-03-19 · Zhan Jin, Yu Luo, Yizhou Zhang, Ziyang Cui 외 arxiv

Conventional pixel-wise loss functions fail to enforce topological constraints in coronary vessel segmentation, producing fragmented vascular trees despite high pixel-level accuracy. We present ARIADNE, a two-stage frame…

MemAudit: Post-hoc Auditing of Poisoned Agent Memory via Causal Attribution and Structural Anomaly Detection

2026-05-22 · Zhewen Tan, Yilun Yao, Huiyan Jin, Wenhan Yu 외 arxiv

Large language model agents increasingly rely on persistent memory to store past interactions, retrieve relevant demonstrations, and improve long-horizon task execution. However, this memory mechanism also creates a prac…

Anomaly Detection

AriadneMem: Threading the Maze of Lifelong Memory for LLM Agents

2026-02-05 · Wenhui Zhu, Xiwen Chen, Zhipeng Wang, Jingjing Wang 외 arxiv

Long-horizon LLM agents require memory systems that remain accurate under fixed context budgets. However, existing systems struggle with two persistent challenges in long-term dialogue: (i) \textbf{disconnected evidence}…