paper-with-me

홈 › Papers

R2-Router: A New Paradigm for LLM Routing with Reasoning

2026-02-02 · Jiaqi Xue, Qian Lou, Jiarong Xing, Heng Huang arxiv

As LLMs proliferate with diverse capabilities and costs, LLM routing has emerged by learning to predict each LLM's quality and cost for a given query, then selecting the one with high quality and low cost. However, existing routers implicitly assume a single fixed quality and cost per LLM for each query, ignoring that the same LLM's quality varies with its output length. This causes routers to exclude powerful LLMs when their estimated cost exceeds the budget, missing the opportunity that these LLMs could still deliver high quality at reduced cost with shorter outputs. To address this, we introduce R2-Router, which treats output length budget as a controllable variable and jointly selects the best LLM and length budget, enforcing the budget via length-constrained instructions. This enables R2-Router to discover that a powerful LLM with constrained output can outperform a weaker LLM at comparable cost-efficient configurations invisible to prior methods. Together with the router framework, we construct R2-Bench, the first routing dataset capturing LLM behavior across diverse output length budgets. Experiments show that R2-Router achieves state-of-the-art performance at 4-5\times lower cost compared with existing routers. This work opens a new direction: routing as reasoning, where routers evolve from reactive selectors to deliberate reasoners that explore which LLM to use and at what cost budget. The code is publicly available at https://github.com/UCF-ML-Research/R2-Router.

📄 PDF Abstract BibTeX arXiv:2602.02823

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RouterEval: A Comprehensive Benchmark for Routing LLMs to Explore Model-level Scaling Up in LLMs

2025-03-08 · Zhongzhan Huang, Guoming Ling, Yupei Lin, Yandong Chen 외

Routing large language models (LLMs) is a new paradigm that uses a router to recommend the best LLM from a pool of candidates for a given input. In this paper, our comprehensive analysis with more than 8,500 LLMs reveals…

Instruction FollowingMathematical Reasoning

GeoRouter: Dynamic Paradigm Routing for Worldwide Image Geolocalization

2026-03-25 · Pengyue Jia, Derong Xu, Yingyi Zhang, Xiaopeng Li 외 arxiv

Worldwide image geolocalization aims to predict precise GPS coordinates for images captured anywhere on Earth, which is challenging due to the large visual and geographic diversity. Recent methods mainly follow two parad…

Select-then-Solve: Paradigm Routing as Inference-Time Optimization for LLM Agents

2026-04-08 · Heng Zhou, Zelin Tan, Zhemeng Zhang, Yutao Fan 외 arxiv

When an LLM-based agent improves on a task, is the gain from the model itself or from the reasoning paradigm wrapped around it? We study this question by comparing six inference-time paradigms, namely Direct, CoT, ReAct,…

TCAndon-Router: Adaptive Reasoning Router for Multi-Agent Collaboration

2026-01-08 · Jiuzhou Zhao, Chunrong Chen, Chenqi Qiao, Lebin Zheng 외 arxiv

Multi-Agent Systems(MAS) have become a powerful paradigm for building high performance intelligent applications. Within these systems, the router responsible for determining which expert agents should handle a given quer…

Reward Model Routing in Alignment

2025-10-03 · Xinle Wu, Yao Lu arxiv

Reinforcement learning from human or AI feedback (RLHF / RLAIF) has become the standard paradigm for aligning large language models (LLMs). However, most pipelines rely on a single reward model (RM), limiting alignment q…

Reinforcement Learning