paper-with-me

Papers

JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety

2026-07-22 · Yuan Xiong, Linji Hao, Shizhu He, Yequan Wang, Lijun Li arxiv

Agent safety is moving from content moderation toward preventing operational failures before tool-using agents act. We propose Janus, a foresight-oriented framework for long-horizon agent safety that trains guards to anticipate delayed risks from partial trajectories. Janus synthesizes diverse agent trajectories via multi-agent simulation and learns a shared policy with two coupled tasks: an anticipation task that forecasts safety-relevant futures and an adjudication task that decides safety from both the observed prefix and anticipated future. The two tasks are jointly optimized with CoAA-RL, which rewards forecasts by their utility for downstream safety judgment. The resulting guard model, Vanguard, blocks unsafe actions before execution. Across four agent-safety benchmarks, Vanguard improves average protection by 15.9 percentage points over baseline guards while increasing benign task completion by 5.1 percentage points.

📄 PDF Abstract BibTeX arXiv:2607.19913

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

JANUS: A Difference-Oriented Analyzer For Financial Centralization Risks in Smart Contracts

2024-12-05 · Wansen Wang, Pu Zhang, Renjie Ji, Wenchao Huang 외

Some smart contracts violate decentralization principles by defining privileged accounts that manage other users' assets without permission, introducing centralization risks that have caused financial losses. Existing me…

TRACE: Trajectory Risk-Aware Compression for Long-Horizon Agent Safety

2026-05-30 · Zhepei Hong, Lin Wang, Liting Li, Haokai Ma 외 arxiv

Long-horizon LLM agents produce safety evidence across long trajectories, where sparse, delayed, and compositional risk signals often escape local moderation. Existing turn-level or short-context detectors struggle to re…

Decoy-Calibrated Failure Audits for Language Models

2026-06-08 · Vyzantinos Repantis, Ameya Gawde, Harshvardhan Singh arxiv

Useful audits reveal not only how often a model fails, but also where its failures concentrate. An auditor may test many candidate explanations: long inputs, indirect questions, distracting evidence, or combinations of t…

A Unified Generative-Predictive Framework for Deterministic Inverse Design

2025-12-10 · Reza T. Batley, Sourav Saha arxiv

Inverse design of heterogeneous material microstructures is a fundamentally ill-posed and famously computationally expensive problem. This is exacerbated by the high-dimensional design spaces associated with finely resol…

The Janus Interface: How Fine-Tuning in Large Language Models Amplifies the Privacy Risks

2023-10-24 · Xiaoyi Chen, Siyuan Tang, Rui Zhu, Shijun Yan 외

The rapid advancements of large language models (LLMs) have raised public concerns about the privacy leakage of personally identifiable information (PII) within their extensive training datasets. Recent studies have demo…

In-Context Learning