paper-with-me

Papers

Probe-Based Interventions for Modifying Agent Behavior

2022-01-26 · Mycal Tucker, William Kuhl, Khizer Shahid, Seth Karten, Katia Sycara, Julie Shah

Neural nets are powerful function approximators, but the behavior of a given neural net, once trained, cannot be easily modified. We wish, however, for people to be able to influence neural agents' actions despite the agents never training with humans, which we formalize as a human-assisted decision-making problem. Inspired by prior art initially developed for model explainability, we develop a method for updating representations in pre-trained neural nets according to externally-specified properties. In experiments, we show how our method may be used to improve human-agent team performance for a variety of neural networks from image classifiers to agents in multi-agent reinforcement learning settings.

📄 PDF Abstract BibTeX arXiv:2201.12938

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingMulti-agent Reinforcement Learningreinforcement-learningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Uncovering Systemic and Environment Errors in Autonomous Systems Using Differential Testing

2025-07-05 · Yashwanthi Anand, Rahil P Mehta, Manish Motwani, Sandhya Saisubramanian arxiv

When an autonomous agent behaves undesirably, including failure to complete a task, it can be difficult to determine whether the behavior is due to a systemic agent error, such as flaws in the model or policy, or an envi…

Inference Time Causal Probing in LLMs

2026-05-08 · Sadegh Khorasani, Saber Salehkaleybar, Negar Kiyavash, Matthias Grossglauser arxiv

Causal probing methods aim to test and control how internal representations influence the behavior of generative models. In causal probing, an intervention modifies hidden states so that a property takes on a different v…

Initial Risk Probing and Feasibility Testing of Glow: a Generative AI-Powered Dialectical Behavior Therapy Skills Coach for Substance Use Recovery and HIV Prevention

2026-02-08 · Liying Wang, Madison Lee, Yunzhang Jiang, Steven Chen 외 arxiv

Background: HIV and substance use represent interacting epidemics with shared psychological drivers - impulsivity and maladaptive coping. Dialectical behavior therapy (DBT) targets these mechanisms but faces scalability …

Learning Explicit Behavioral Models with Adaptive Questions and World-Model Probes

2026-06-05 · Hikaru Shindo, Yu Deng, Teng Cao, Quentin Delfosse 외 arxiv

Interactive agents trained only against task return can achieve high scores while failing to represent the mechanisms that make their actions succeed. This makes brittle behavior difficult to diagnose and limits adaptati…

Question AnsweringCode Repair

Automated Attribution Graph Interpretation via Probe Prompting

2025-11-10 · Giuseppe Birardi, Gonçalo Paulo arxiv

Even though we know the precise computations that lead from a large language model (LLM) input to its output this computation remains very hard to interpret. One way to make it easier to understand this process is by cre…