paper-with-me

홈 › Papers

Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards

2026-07-24 · Ruoxi Cheng, Haoxuan Ma, Hongyi Zhang, Junming Zhang, Ranjie Duan, Qiaolin Xia, Hao Wang, Yu Lu, Haibo Shi, Xingjun Ma arxiv

Search-augmented language agents should retrieve external information only when necessary and ground their answers in retrieved evidence. Existing external rewards provide either sparse outcome supervision or richer feedback from process annotations and LLM judges. Outcome rewards scale readily but cannot distinguish grounded retrieval from redundant search, whereas richer signals require costly annotation or inference during training. Internal rewards based on policy-side signals such as entropy, likelihood, or information gain are graded and inexpensive to evaluate, yet mainly reflect model confidence rather than evidence grounding. We propose Search-G1, a representation-based intrinsic reward framework that measures the operational grounding of an agent's answers through two intervention-calibrated readouts. A prompt-state readout predicts closed-book sufficiency, whose complement defines policy-relative retrieval necessity; an answer-commit readout estimates evidence reliance from answer-stage sensitivity to evidence deletion. Together, they provide additional credit to correct searched trajectories when retrieval is estimated necessary and the answer is evidence-sensitive, favor correct direct answers when closed-book knowledge suffices, and penalize repeated search. After calibration, reward scoring requires neither process annotations nor LLM-as-judge inference during policy optimization. Because reinforcement learning changes policy representations, Search-G1 periodically refits both readouts on trajectories from the latest checkpoint, allowing the reward to co-evolve with the policy. Experiments across multiple search-based question-answering benchmarks and two model scales show that Search-G1 improves the grounding--search-cost trade-off, producing shorter response-side trajectories at competitive task accuracy. Code is available at https://github.com/Rosy0912/Search-G1.

📄 PDF Abstract BibTeX arXiv:2608.07531

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LiveBrowseComp: Are Search Agents Searching, or Just Verifying What They Already Know?

2026-05-27 · HuiMing Fan, Xiao Wang, Zheng Chu, Qianyu Wang 외 arxiv

Are LLM-based search agents genuinely searching, or using the web to verify what they already know? We study this question on BrowseComp with three diagnostics. Our analysis reveals Intrinsic Knowledge Dependence (IKD): …

Deep Reinforcement Learning for Multi-Agent Interaction

2022-08-02 · Ibrahim H. Ahmed, Cillian Brewitt, Ignacio Carlucho, Filippos Christianos 외

The development of autonomous agents which can interact with other agents to accomplish a given task is a core area of research in artificial intelligence and machine learning. Towards this goal, the Autonomous Agents Re…

BIG-bench Machine LearningCausal InferenceDeep Reinforcement LearningMulti-agent Reinforcement Learning+4

Emotion in Reinforcement Learning Agents and Robots: A Survey

2017-05-15 · Thomas M. Moerland, Joost Broekens, Catholijn M. Jonker

This article provides the first survey of computational models of emotion in reinforcement learning (RL) agents. The survey focuses on agent/robot emotions, and mostly ignores human user emotions. Emotions are recognized…

AI AgentDecision Makingreinforcement-learningReinforcement Learning+2

PreND: Enhancing Intrinsic Motivation in Reinforcement Learning through Pre-trained Network Distillation

2024-10-02 · Mohammadamin Davoodabadi, Negin Hashemi Dijujin, Mahdieh Soleymani Baghshah

Intrinsic motivation, inspired by the psychology of developmental learning in infants, stimulates exploration in agents without relying solely on sparse external rewards. Existing methods in reinforcement learning like R…

Developmental Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

LLM-enabled Social Agents

2026-05-04 · Önder Gürcan, Moharram Challenger arxiv

Large Language Models (LLMs) have transformed agent-agent and human-agent interaction by enabling software, physical, and simulation agents to communicate and deliberate through natural language. Yet fluent language use …