RICE-PO: Turning Retrieval Interactions into Credit Signals for Reasoning Agents
Retrieval is increasingly moving from one-shot matching toward interactive reasoning, where language agents iteratively inspect evidence, reformulate queries, and search again. Training such agents raises a credit-assignment challenge: executable actions such as queries or summaries can be directly evaluated by the retriever, while latent reasoning steps are not directly observable and only affect future executable actions. This asymmetry makes outcome-level reward assignment unreliable, as the same final reward may credit reasoning steps that did not actually shape retrieval success. We propose RICE-PO, a critic-free policy optimization framework that converts retrieval interactions into localized learning signals. RICE-PO selects high-uncertainty executable actions as anchors, evaluates local counterfactual branches using retrieval metrics, and propagates credit to latent reasoning steps only when reasoning-to-action influence is strong and future residual effects are stable. On BRIGHT and BEIR, RICE-PO consistently outperforms prompt-based agents and group-based RL baselines under the same retriever setting. These results show that the structure of agent-environment interaction itself can provide useful supervision for training reasoning-based retrieval agents.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards
Training deep research agents, namely systems that plan, search, evaluate evidence, and synthesize long-form reports, pushes reinforcement learning beyond the regime of verifiable rewards. Their outputs lack ground-truth…
Reinforcement LearningPropagation of a carbon price in a credit portfolio through macroeconomic factors
We study how the climate transition through a low-carbon economy, implemented by carbon pricing, propagates in a credit portfolio and precisely describe how carbon price dynamics affects credit risk measures such as prob…
A New Model for Pricing Collateralized Financial Derivatives
This paper presents a new model for pricing financial derivatives subject to collateralization. It allows for collateral arrangements adhering to bankruptcy laws. As such, the model can back out the market price of a col…
modelAnalytical Pricing of 2 Factor Structural PDE model for a Puttable Bond with Credit Risk
In this paper is proposed a 2 factor structural PDE model of pricing puttable bond with credit risk and derived the analytical pricing formula. To this end, first, a 2 factor structural (PDE) model of pricing zero coupon…
A Credit Assignment Compiler for Joint Prediction
Many machine learning applications involve jointly predicting multiple mutually dependent output variables. Learning to search is a family of methods where the complex decision problem is cast into a sequence of decision…
Prediction