paper-with-me

홈 › Papers

BEAR: Budgeted Evidence Allocation for Multi-Document Reasoning

2026-01-26 · Lin Sun, Linglin Zhang, Jingang Huang, Change Jia, Zhengwei Cheng, Xiangzheng Zhang arxiv

We argue that multi-document reasoning is constrained not only by how much text a model can read, but also by how limited query-time evidence budget is allocated across documents and semantic granularities. Full-context inference exposes the model to broad evidence non-selectively and at high per-query cost, while flat chunk retrieval often returns locally relevant passages that are weakly organized for cross-document synthesis. We present \textbf{BEAR}, a framework for structured evidence allocation that builds hierarchical semantic indices offline and performs coarse-to-fine evidence access at query time through complementary \emph{exploration} and \emph{recovery} paths. This coarse-to-fine design can be viewed as structured evidence allocation under a fixed evidence-context budget. Across synthetic and real-world benchmarks, BEAR performs particularly strongly on DragonBall, remains competitive with strong retrieval-based baselines on HotpotQA, and yields the best retrieval-based result on 2Wiki under our evaluated protocol, while operating under substantially smaller \emph{query-time evidence budgets} than the reported long-context references. Additional analyses suggest that the gains are associated with hierarchy as an allocation substrate together with complementary exploration and recovery, rather than semantic chunking alone.

📄 PDF Abstract BibTeX arXiv:2601.18116

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Authority Expectancy Effect in Multi-Party Conflict

2026-08-08 · Eunna Lee arxiv

In multi-party competitive settings, across experiments on resource allocation, fault attribution, and dispute mediation, we test whether social authority (SA) cues such as occupational status and institutional documenta…

DSA: More Efficient Budgeted Pruning via Differentiable Sparsity Allocation

2020-04-05 · ECCV 2020 8 · Xuefei Ning, Tianchen Zhao, Wenshuo Li, Peng Lei 외

Budgeted pruning is the problem of pruning under resource constraints. In budgeted pruning, how to distribute the resources across layers (i.e., sparsity allocation) is the key problem. Traditional methods solve it by di…

VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding

2026-05-07 · Kuanwei Lin, Wenhao Zhang, Ge Li arxiv

Video large multimodal models increasingly face a scalability bottleneck: long videos produce excessively long visual-token sequences, which sharply increase memory and latency during inference. While existing compressio…

Budgeted Attention Allocation: Cost-Conditioned Compute Control for Efficient Transformers

2026-05-07 · Amrit Nidhi arxiv

Transformers usually expose one inference cost per trained model, while deployed systems often need multiple cost-quality operating points. We study Budgeted Attention Allocation, a monotone head-gating mechanism conditi…

EMBER: Efficient Memory via Budgeted Evidence Retention for Long-Horizon Agents

2026-06-04 · Yilong Li, Suman Banerjee, Tong Che arxiv

Long-horizon agents can archive large histories, but future answers still incur retrieval, rereading, and context costs. When retained memory misses answer-relevant evidence, the system must return to larger portions of …