paper-with-me

홈 › Papers

Context-Picker: Dynamic context selection using multi-stage reinforcement learning

2025-12-16 · Siyuan Zhu, Chengdong Xu, Kaiqiang Ke, Chao Yu arxiv

In long-context question answering, selecting the appropriate scope of context for a query remains a key and unresolved challenge. Insufficient context can lead to missing essential information, whereas excessive context often introduces noise and degrades answer quality. Conventional methods, such as retrieving a fixed number of passages or applying reranking, struggle to dynamically determine which context to include. This is especially problematic for factoid questions, which typically depend only on a few precise pieces of evidence. To overcome this limitation, we propose Context-Picker, a reasoning-aware framework that reframes context selection as the task of identifying a minimal sufficient evidence subset, moving beyond conventional similarity-based ranking. Context-Picker uses a human-inspired two-stage reinforcement learning schedule: stage 1 focuses on improving the recall rate of critical passages, and stage 2 prioritizes pruning redundancy to distill a compact evidence set. To resolve reward sparsity, we propose an offline evidence distillation pipeline that mines ``minimal sufficient sets" via a Leave-One-Out (LOO) procedure, providing dense and task-aligned supervision. Experiments on five long-context and multi-hop QA datasets demonstrate that our method outperforms strong RAG baselines and achieved higher answer accuracy. Ablation studies also indicate that our coarse-to-fine optimization schedule, the redundancy-aware reward shaping, along with the rationale generated by the policy, all contribute substantially to these gains.

📄 PDF Abstract BibTeX arXiv:2512.14465

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningQuestion Answering

Similar Papers 제목 키워드 기반

Context-Aware Synthesis of Optimization Pipelines for Warehouse Optimization

2026-06-25 · Janik Bischoff, Anne Meyer, Uta Mohring, Fabian Dunke 외 arxiv

Order fulfillment in manual picker-to-goods warehouses involves interconnected decisions such as item assignment, order batching, and picker routing. While integrated models capture interactions between these decisions, …

Learning to Solve the Min-Max Mixed-Shelves Picker-Routing Problem via Hierarchical and Parallel Decoding

2025-02-14 · Laurin Luttmann, Lin Xie

The Mixed-Shelves Picker Routing Problem (MSPRP) is a fundamental challenge in warehouse logistics, where pickers must navigate a mixed-shelves environment to retrieve SKUs efficiently. Traditional heuristics and optimiz…

Decision MakingMulti-agent Reinforcement LearningNavigateSequential Decision Making

Enhance Incomplete Utterance Restoration by Joint Learning Token Extraction and Text Generation

2022-04-08 · NAACL 2022 7 · Shumpei Inoue, Tsungwei Liu, Nguyen Hong Son, Minh-Tien Nguyen

This paper introduces a model for incomplete utterance restoration (IUR) called JET (\textbf{J}oint learning token \textbf{E}xtraction and \textbf{T}ext generation). Different from prior studies that only work on extract…

Language ModelingLanguage ModellingText Generation

AdvPicker: Effectively Leveraging Unlabeled Data via Adversarial Discriminator for Cross-Lingual NER

2021-06-04 · ACL 2021 5 · WEILE CHEN, Huiqiang Jiang, Qianhui Wu, Börje F. Karlsson 외

Neural methods have been shown to achieve high performance in Named Entity Recognition (NER), but rely on costly high-quality labeled data for training, which is not always available across languages. While previous work…

Cross-Lingual NERMachine Translationnamed-entity-recognitionNamed Entity Recognition+3

Breakout-picker: Reducing false positives in deep learning-based borehole breakout characterization from acoustic image logs

2026-04-17 · Guangyu Wang, Xiaodong Ma, Xinming Wu arxiv

Borehole breakouts are stress-induced spalling on the borehole wall, which are identifiable in acoustic image logs as paired zones with near-symmetry azimuths, low acoustic amplitudes, and increased borehole radius. Accu…