paper-with-me

홈 › Papers

Towards Trustworthy Multi-Turn LLM Agents via Behavioral Guidance

2025-12-12 · Gonca Gürsun arxiv

Large Language Models demonstrate strong reasoning and generation abilities, yet their behavior in multi-turn tasks often lacks reliability and verifiability. We present a task completion framework that enables LLM-based agents to act under explicit behavioral guidance in environments described by reinforcement learning formalisms with defined observation, action, and reward signals. The framework integrates three components: a lightweight task profiler that selects reasoning and generation strategies, a reasoning module that learns verifiable observation - action mappings, and a generation module that enforces constraint-compliant outputs through validation or deterministic synthesis. We show that as the agent interacts with the environment, these components co-evolve, yielding trustworthy behavior.

📄 PDF Abstract BibTeX arXiv:2512.11421

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

NetAgentBench: A State-Centric Benchmark for Evaluating Agentic Network Configuration

2026-04-03 · Ahmed Twabi, Yepeng Ding, Tohru Kondo arxiv

As agentic network management gains popularity, there is a critical need for evaluation frameworks that transcend static, one-shot testing. To address this, we introduce NetAgentBench, a dynamic benchmark that evaluates …

Designing The Internet of Agents: A Framework for Trustworthy, Transparent, and Collaborative Human-Agent Interaction (HAX)

2025-12-12 · Marc Scibelli, Krystelle Gonzalez Papaux, Julia Valenti, Srishti Kush arxiv

The rise of generative and autonomous agents marks a fundamental shift in computing, demanding a rethinking of how humans collaborate with probabilistic, partially autonomous systems. We present the Human-AI-Experience (…

What Makes Agent Memory Useful for Reliable Unanswerable Question Handling?

2026-08-28 · Chuanyuan Tan, Junjie Yu, Yuxin Wang, Yining Zheng 외 arxiv

Reliable handling of unanswerable questions (UAQs) is critical for trustworthy LLM-based agents. Although memory is widely used in agent systems, its role in reliable UAQ handling remains unclear. We present a systematic…

Toddler-Guidance Learning: Impacts of Critical Period on Multimodal AI Agents

2022-01-12 · Junseok Park, Kwanyoung Park, Hyunseok Oh, Ganghun Lee 외

Critical periods are phases during which a toddler's brain develops in spurts. To promote children's cognitive development, proper guidance is critical in this stage. However, it is not clear whether such a critical peri…

Reinforcement Learning (RL)Transfer Learning

Behavioral Determinants of Deployed AI Agents in Social Networks: A Multi-Factor Study of Personality, Model, and Guardrail Specification

2026-05-08 · Sarah Wilson, Diem Linh Dang, Usman Ali Moazzam, Shan Ye 외 arxiv

Autonomous AI agents are increasingly deployed in open social environments, yet the relationship between their configuration specifications and their emergent social behavior remains poorly understood. We present a contr…