paper-with-me

홈 › Papers

TRACER: Trajectory Risk Aggregation for Critical Episodes in Agentic Reasoning

2026-02-11 · Sina Tayebati, Divake Kumar, Nastaran Darabi, Davide Ettori, Ranganath Krishnan, Amit Ranjan Trivedi arxiv

Estimating uncertainty for AI agents in real-world multi-turn tool-using interaction with humans is difficult because failures are often triggered by sparse critical episodes (e.g., looping, incoherent tool use, or user-agent miscoordination) even when local generation appears confident. Existing uncertainty proxies focus on single-shot text generation and therefore miss these trajectory-level breakdown signals. We introduce TRACER, a trajectory-level uncertainty metric for dual-control Tool-Agent-User interaction. TRACER combines content-aware surprisal with situational-awareness signals, semantic and lexical repetition, and tool-grounded coherence gaps, and aggregates them using a tail-focused risk functional with a MAX-composite step risk to surface decisive anomalies. We evaluate TRACER on $τ^2$-bench by predicting task failure and selective task execution. To this end, TRACER improves AUROC by up to 37.1% and AUARC by up to 55% over baselines, enabling earlier and more accurate detection of uncertainty in complex conversational tool-use settings. Our code and benchmark are available at https://github.com/sinatayebati/agent-tracer.

📄 PDF Abstract BibTeX arXiv:2602.11409

Code (0)

등록된 구현이 없습니다.

Tasks

Text Generation

Similar Papers 제목 키워드 기반

SLIM-RL: Risk-Budgeted Random-Masking RL for Diffusion LLMs Without Trajectory Slicing

2026-06-30 · Ruikang Zhao, Zhenting Wang, Han Gao, Ligong Han arxiv

Reinforcement learning for diffusion large language models (dLLMs) has largely moved to trajectory-aware methods. The current state of the art, TraceRL, holds that random masking is mismatched with the model's inference …

Reinforcement Learning

HalluTracer: Hallucination Detection via Depth-Averaging Truth Signals

2026-08-17 · Zhihao Guo, Zonghan Wu, Huan Huo, DaYong Ye 외 arxiv

Even well-aligned large language models confidently generate factually incorrect text, making hallucination a persistent reliability risk in high-stakes deployments. These models nonetheless carry linearly separable trut…

TraCeR: Transformer-Based Competing Risk Analysis with Longitudinal Covariates

2025-12-19 · Maxmillan Ries, Sohan Seth arxiv

Survival analysis is a critical tool for modeling time-to-event data. Recent deep learning-based models have reduced various modeling assumptions including proportional hazard and linearity. However, a persistent challen…

TRACER: A Framework for Facilitating Accurate and Interpretable Analytics for High Stakes Applications

2020-03-24 · Kaiping Zheng, Shaofeng Cai, Horng Ruey Chua, Wei Wang 외

In high stakes applications such as healthcare and finance analytics, the interpretability of predictive models is required and necessary for domain practitioners to trust the predictions. Traditional machine learning mo…

Feature ImportanceManagementTime SeriesTime Series Analysis

TRACER: Transfer Learning based Real-time Adaptation for Clinical Evolving Risk

2025-12-14 · Mengying Yan, Ziye Tian, Siqi Li, Nan Liu 외 arxiv

Clinical decision support tools built on electronic health records often experience performance drift due to temporal population shifts, particularly when changes in the clinical environment initially affect only a subse…

Transfer Learning