paper-with-me

홈 › Papers

Don't Start What You Can't Finish: A Counterfactual Audit of Support-State Triage in LLM Agents

2026-04-17 · Eren Unlu arxiv

Current agent evaluations largely reward execution on fully specified tasks, while recent work studies clarification [11, 22, 2], capability awareness [9, 1], abstention [8, 14], and search termination [20, 5] mostly in isolation. This leaves open whether agents can diagnose why a task is blocked before acting. We introduce the Support-State Triage Audit (SSTA-32), a matched-item diagnostic framework in which minimal counterfactual edits flip the same base request across four support states: Complete (ANSWER), Clarifiable (CLARIFY), Support-Blocked (REQUEST SUPPORT), and Unsupported-Now (ABSTAIN). We evaluate a frontier model under four prompting conditions - Direct, Action-Only, Confidence-Only, and a typed Preflight Support Check (PSC) - using Dual-Persona Auto-Auditing (DPAA) with deterministic heuristic scoring. Default execution overcommits heavily on non-complete tasks (41.7% overcommitment rate). Scalar confidence mapping avoids overcommitment but collapses the three-way deferral space (58.3% typed deferral accuracy). Conversely, both Action-Only and PSC achieve 91.7% typed deferral accuracy by surfacing the categorical ontology in the prompt. Targeted ablations confirm that removing the support-sufficiency dimension selectively degrades REQUEST SUPPORT accuracy, while removing the evidence-sufficiency dimension triggers systematic overcommitment on unsupported items. Because DPAA operates within a single context window, these results represent upper-bound capability estimates; nonetheless, the structural findings indicate that frontier models possess strong latent triage capabilities that require explicit categorical decision paths to activate safely.

📄 PDF Abstract BibTeX arXiv:2604.16752

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A New Paradigm for Counterfactual Reasoning in Fairness and Recourse

2024-01-25 · Lucius E. J. Bynum, Joshua R. Loftus, Julia Stoyanovich

Counterfactuals and counterfactual reasoning underpin numerous techniques for auditing and understanding artificial intelligence (AI) systems. The traditional paradigm for counterfactual reasoning in this literature is t…

counterfactualCounterfactual ReasoningFairness

What If They Took the Shot? A Hierarchical Bayesian Framework for Counterfactual Expected Goals

2025-11-28 · Mikayil Mahmudlu, Oktay Karakuş, Hasan Arkadaş arxiv

This study develops a hierarchical Bayesian framework that integrates expert domain knowledge to quantify player-specific effects in expected goals (xG) estimation, addressing a limitation of standard models that treat a…

EviSnap: Faithful Evidence-Cited Explanations for Cold-Start Cross-Domain Recommendation

2026-01-09 · Yingjun Dai, Ahmed El-Roby arxiv

Cold-start cross-domain recommender (CDR) systems predict a user's preferences in a target domain using only their source-domain behavior, yet existing CDR models either map opaque embeddings or rely on post-hoc or LLM-g…

All Required, In Order: Phase-Level Evaluation for AI-Human Dialogue in Healthcare and Beyond

2026-01-13 · Shubham Kulkarni, Alexander Lyzhov, Shiva Chaitanya, Preetam Joshi arxiv

Conversational AI is starting to support real clinical work, but most evaluation methods miss how compliance depends on the full course of a conversation. We introduce Obligatory-Information Phase Structured Compliance E…

Theory Under Construction: Orchestrating Language Models for Research Software Where the Specification Evolves

2026-04-29 · Halley Young, Nikolaj Björner arxiv

Large language models can now generate substantial code and draft research text, but research-software projects require more than either artifact alone. The mathematical thesis, executable system, benchmark surface, and …