paper-with-me

홈 › Papers

Evaluating Generative Agents with Actions Grounded in Socially Distributed Task Environments using Incognita

2026-07-03 · Dan C. Hsu, Luke Lu arxiv

Effective agency in social environments depends on when an agent seeks knowledge, when it acts, and whether its actions are justified by acquired information. Existing grounded benchmarks provide executable actions, persistent state, and verifiable outcomes, while social simulation environments provide rich interaction among language agents. We study an evaluation setting that combines these requirements. We define socially distributed task environments as interactive environments where task-relevant knowledge is partitioned across role-isolated participants and consequential actions are accessible only through them. Communication serves as exploration over role-partitioned knowledge, while grounded action serves as exploitation over environment state. We introduce Incognita, a Concordia-based framework that separates social interaction from grounded execution. The evaluated agent routes messages to a user or specialist entities; specialists mediate admissible operations; a deterministic sub-environment executes accepted operations over a canonical state; and an offline evaluator scores outcomes with inherited rewards. Incognita-Retail transforms tau-bench retail into a multi-entity environment while preserving final-state reward semantics. We evaluate three generative agent models on 18 tasks stratified by social breadth, with 540 trials. Progress appears in reward and behavior: success rises from 0 percent to 8.9 percent and 17.2 percent, while premature finalization falls from 100 percent to 87 percent and 58 percent. Stronger models elicit more hidden knowledge, contact more entities, and attempt more grounded writes, yet reliability remains low. These findings show that socially distributed task environments expose behavior before reliable success, including knowledge elicitation, source selection, grounded action attempts, and premature completion belief.

📄 PDF Abstract BibTeX arXiv:2607.02975

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DoubleAgents: Human-Agent Alignment in a Socially Embedded Workflow

2025-09-16 · Tao Long, Xuanming Zhang, Sitong Wang, Zhou Yu 외 arxiv

Aligning agentic AI with user intent is critical for delegating complex, socially embedded tasks, yet user preferences are often implicit, evolving, and difficult to specify upfront. We present DoubleAgents, a system for…

Online Agent-as-a-Judge: Situation-Generating Evaluation for Interactive Agents

2026-06-06 · Hyogon Ryu, Jeonghwan Kim, Yewon Lim, Chaeun Lee 외 arxiv

Evaluating LLM-powered interactive social agents is challenging because socially relevant behaviors depend not only on isolated outputs, but also on prior interactions, social roles, and downstream actions. Existing meth…

ProMediate: A Socio-cognitive framework for evaluating proactive agents in multi-party negotiation

2025-10-29 · Ziyi Liu, Bahar Sarrafzadeh, Pei Zhou, Longqi Yang 외 arxiv

While Large Language Models (LLMs) are increasingly used in agentic frameworks to assist individual users, there is a growing need for agents that can proactively manage complex, multi-party collaboration. Systematic eva…

A Multimodal Framework for Human-Multi-Agent Interaction

2026-03-24 · Shaid Hasan, Breenice Lee, Sujan Sarker, Tariq Iqbal arxiv

Human-robot interaction is increasingly moving toward multi-robot, socially grounded environments. Existing systems struggle to integrate multimodal perception, embodied expression, and coordinated decision-making in a u…

Multimodal Reasoning

Sleepless Nights, Sugary Days: Creating Synthetic Users with Health Conditions for Realistic Coaching Agent Interactions

2025-02-18 · Taedong Yun, Eric Yang, Mustafa Safdari, Jong Ha Lee 외

We present an end-to-end framework for generating synthetic users for evaluating interactive agents designed to encourage positive behavior changes, such as in health and lifestyle coaching. The synthetic users are groun…