paper-with-me

홈 › Papers

RASER: Recoverability-Aware Selective Escalation Router for Multi-Hop Question Answering

2026-06-01 · Yuyang Li, Zihe Yan, Tobias Käfer arxiv

Multi-hop question-answering systems often use expensive retrieval on every question. They may decompose the question, run several retrieval rounds, or search through bridge entities before answering. All of these strategies rely on repeated LLM calls to rewrite or decompose the question, which increases extra token cost, and it is not fitting when the LLM budget is tight. However, our analysis shows that lots of multi-hop questions are already answered correctly by a single one-shot RAG, so running an extra retrieval on every question wastes the budget. We introduce RASER (Recoverability-Aware Selective Escalation Router), a family of cheap routers built on one-shot RAG and six features from it. RASER-2 decides whether to stop or escalate to the extra-retrieval action PRUNE. RASER-3 chooses among one-shot RAG, PRUNE, and iterative retrieval IRCoT, using the same features but adding an explicit cost-accuracy trade-off. Neither router makes an extra LLM call to decide. Across six LLMs and three multi-hop QA benchmarks, both routers stay competitive with the other state-of-the-art (SOTA) baselines in F1 while spending only 41-49% of always-prune's tokens and also less than the iterative and decomposition retrieval baselines.

📄 PDF Abstract BibTeX arXiv:2606.02488

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-hop Question Answering

Similar Papers 제목 키워드 기반

Modality Relevance is not Modality Utility: Post-hoc Selective Modality Escalation for Cost-Aware Multimodal RAG

2026-07-03 · Xue Li, Yiming Gai arxiv

Multimodal retrieval-augmented generation (RAG) grounds a generator in evidence drawn from heterogeneous modalities -- text, tables, and images. The dominant deployment choice is binary and made before the model has trie…

CascadeDebate: Multi-Agent Deliberation for Cost-Aware LLM Cascades

2026-04-14 · Raeyoung Chang, Dongwook Kwon, Jisoo Lee, Nikhil Verma arxiv

Cascaded LLM systems coordinate models of varying sizes with human experts to balance accuracy, cost, and abstention under uncertainty. However, single-model tiers at each stage often struggle with ambiguous queries, tri…

General Knowledge

Expected Gain-based Escalation in Vertical Federated Learning

2026-06-30 · Mohamad Mestoukirdi, Vincent Corlay arxiv

Collaborative inference can improve predictive performance by integrating complementary information across agents, but applying collaborative fusion to every sample can incur unnecessary communication and computational o…

Federated Learning

SAFE-Cascade: Cost-Adaptive Vision-Language Routing for Chart Question Answering

2026-06-17 · Ayush Dwivedi, Qixin Wang, Ashvi Soni, Ruoteng Wang 외 arxiv

Vision-language models (VLMs) are powerful for chart question answering, but invoking a VLM for every query can be unnecessarily expensive when many questions are answerable from OCR text and lightweight language reasoni…

Chart Question AnsweringVisual Grounding

xRouter: Training Cost-Aware LLMs Orchestration System via Reinforcement Learning

2025-10-09 · Cheng Qian, Zuxin Liu, Shirley Kokane, Akshara Prabhakar 외 arxiv

Modern LLM deployments confront a widening cost-performance spectrum: premium models deliver strong reasoning but are expensive, while lightweight models are economical yet brittle on complex tasks. Static escalation rul…

Reinforcement Learning