paper-with-me

홈 › Papers

Governance by Construction for Generalist Agents

2026-05-20 · Segev Shlomov, Iftach Shoham, Alon Oved, Ido Levy, Sami Marreed, Harold Ship, Offer Akrabi, Sergey Zeltyn, Avi Yaeli, Nir Mashkif arxiv

Enterprise agents are increasingly expected to operate autonomously across tools and interfaces, yet production deployments require governance by construction. Systems must specify which actions are allowed, when human oversight is required, and what information may be exposed, without rebuilding the agent for each domain. This demo presents CUGA's policy system, a modular policy-as-code layer that composes with a generalist LLM agent to deliver predictable, auditable, and compliance-aware behavior in compound workflows without model fine-tuning. We present a runtime governance architecture that enforces policy interventions at every critical stage of execution. Rather than passively constraining behavior, policies intercept the agent at five structural checkpoints: upstream of planning (Intent Guard), within the system prompt to steer reasoning (Playbook), at the tool-call boundary to enforce proper usage (Tool Guide), outside the reasoning loop as a Human-in-the-Loop gate for high-risk actions (Tool Approvals), and at the output stage to filter and structure the final response (Output Formatter). Together, these stages embed governance continuously across the agent's execution pipeline rather than treating it as an afterthought. Using a healthcare scenario and a multi-layered enforcement intervention, the demo shows dynamic playbook injection for structured tool-sequence enforcement, intent guards that block malicious or accidental harmful requests, and human-in-the-loop tool approval checkpoints for potentially destructive actions. The artifact illustrates how typed governance primitives enable faster, safer deployment of enterprise agentic systems while improving policy adherence and execution consistency.

📄 PDF Abstract BibTeX arXiv:2605.20874

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Artificial collectives of specialists and generalists excel at different tasks

2026-06-18 · John Meluso, Laurent Hébert-Dufresne, Christoph Riedl, H. Oliver Gao arxiv

Collective artificial intelligence, where multiple agents work on shared tasks, holds potential to solve expansive problems in fields from medicine to collective governance. But while prescriptive engineering solutions a…

From Benchmarks to Business Impact: Deploying IBM Generalist Agent in Enterprise Production

2025-10-27 · Segev Shlomov, Alon Oved, Sami Marreed, Ido Levy 외 arxiv

Agents are rapidly advancing in automating digital work, but enterprises face a harder challenge: moving beyond prototypes to deployed systems that deliver measurable business value. This path is complicated by fragmente…

What Must Generalist Agents Remember?

2026-06-17 · Khurram Yamin, Namrata Deka, Maitreyi Swaroop, Albert Ting 외 arxiv

This paper develops a formal account of what generalist agents must store in memory in order to act near-optimally across multiple environments and goals. It shows that when two domains share an observational bottleneck …

Mind2Web: Towards a Generalist Agent for the Web

2023-06-09 · NeurIPS 2023 11 · Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen 외

We introduce Mind2Web, the first dataset for developing and evaluating generalist agents for the web that can follow language instructions to complete complex tasks on any website. Existing datasets for web agents either…

Initiation Safety: A Missing Dimension in Generalist-Robot Safety

2026-07-08 · Zhijin Meng, Francisco Cruz arxiv

Safety for generalist robots is usually discussed in terms of motion or dialogue. We argue a third question is missing: should the robot take its first hard-to-undo social action at all, such as a greeting, an uninvited …