paper-with-me

Papers

Testing Language Model Agents Safely in the Wild

2023-11-17 · Silen Naihin, David Atkinson, Marc Green, Merwane Hamadi, Craig Swift, Douglas Schonholtz, Adam Tauman Kalai, David Bau

A prerequisite for safe autonomy-in-the-wild is safe testing-in-the-wild. Yet real-world autonomous tests face several unique safety challenges, both due to the possibility of causing harm during a test, as well as the risk of encountering new unsafe agent behavior through interactions with real-world and potentially malicious actors. We propose a framework for conducting safe autonomous agent tests on the open internet: agent actions are audited by a context-sensitive monitor that enforces a stringent safety boundary to stop an unsafe test, with suspect behavior ranked and logged to be examined by humans. We design a basic safety monitor (AgentMonitor) that is flexible enough to monitor existing LLM agents, and, using an adversarial simulated agent, we measure its ability to identify and stop unsafe situations. Then we apply the AgentMonitor on a battery of real-world tests of AutoGPT, and we identify several limitations and challenges that will face the creation of safe in-the-wild tests as autonomous agents grow more capable.

📄 PDF Abstract BibTeX arXiv:2311.10538

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modellingmodel

Similar Papers 제목 키워드 기반

AI Agents for Web Testing: A Case Study in the Wild

2025-09-05 · Naimeng Ye, Xiao Yu, Ruize Xu, Tianyi Peng 외 arxiv

Automated web testing plays a critical role in ensuring high-quality user experiences and delivering business value. Traditional approaches primarily focus on code coverage and load testing, but often fall short of captu…

Do Pedestrians Pay Attention? Eye Contact Detection in the Wild

2021-12-08 · Younes Belkada, Lorenzo Bertoni, Romain Caristan, Taylor Mordan 외

In urban or crowded environments, humans rely on eye contact for fast and efficient communication with nearby people. Autonomous agents also need to detect eye contact to interact with pedestrians and safely navigate aro…

Autonomous VehiclesContact DetectionDomain AdaptationNavigate

RICoTA: Red-teaming of In-the-wild Conversation with Test Attempts

2025-01-29 · Eujeong Choi, Younghun Jeong, SooMin Kim, Won Ik Cho

User interactions with conversational agents (CAs) evolve in the era of heavily guardrailed large language models (LLMs). As users push beyond programmed boundaries to explore and build relationships with these systems, …

ChatbotRed Teaming

From Segments to Scenes: Temporal Understanding for Agentic Autonomous Driving via Vision-Language Models

2025-12-04 · Kevin Cannons, Saeed Ranjbar Alvar, Mohammad Asiful Hossain, Ahmad Rezaei 외 arxiv

Vision-Language Models (VLMs) are increasingly deployed as the perception and reasoning backbone of autonomous agents acting in the wild, with autonomous driving (AD) being one of the most safety-critical instances. Reli…

Autonomous Driving

Wildfire Smoke and Air Quality: How Machine Learning Can Guide Forest Management

2020-10-09 · Lorenzo Tomaselli, Coty Jen, Ann B. Lee

Prescribed burns are currently the most effective method of reducing the risk of widespread wildfires, but a largely missing component in forest management is knowing which fuels one can safely burn to minimize exposure …

BIG-bench Machine LearningClusteringManagement