paper-with-me

홈 › Papers

Unsolvability Ceiling in Multi-LLM Routing: An Empirical Study of Evaluation Artifacts

2026-05-08 · Saloni Garg, Amit Sagtani arxiv

Efficient routing across multiple LLMs enables cost-quality tradeoffs by directing queries to the cheapest capable model. Prior work attributes routing headroom to an "unsolvability ceiling", queries no model in the pool can solve. We present a large-scale study of multi-tier LLM routing with 206,000 query-model pairs across six benchmarks (MMLU, MedQA, HumanEval, MBPP, Alpaca, ShareGPT) using the Gemma 4 and Llama 3.1 families. Evaluating with both LLM-as-a-judge and exact-match metrics, we show that a substantial portion of reported unsolvability stems from evaluation artifacts: (i) systematic judge biases favoring verbosity over correctness, (ii) truncation under fixed generation budgets, and (iii) output format mismatches. Through dual-judge validation and exact-match grounding, we reduce measured unsolvability across tasks. We introduce a decomposition framework attributing failures to these artifacts, revealing consistent patterns across domains and model families. These artifacts also distort router training signals: standard routers collapse to majority-class prediction (~79% smallest-tier optimal), confirmed via random-feature and shuffled-label controls, incurring a 13-17 percentage point opportunity cost. We provide actionable recommendations including dual-judge validation, exact-match anchoring, and cost-sensitive objectives. Our findings suggest existing routing headroom estimates are substantially inflated, underscoring the need for reliable evaluation protocols in multi-LLM systems.

📄 PDF Abstract BibTeX arXiv:2605.07395

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Why Couldn't You do that? Explaining Unsolvability of Classical Planning Problems in the Presence of Plan Advice

2019-03-19 · Sarath Sreedharan, Siddharth Srivastava, David Smith, Subbarao Kambhampati

Explainable planning is widely accepted as a prerequisite for autonomous agents to successfully work with humans. While there has been a lot of research on generating explanations of solutions to planning problems, expla…

Learning the Boundary of Solvability: Aligning LLMs to Detect Unsolvable Problems

2025-12-01 · Dengyun Peng, Qiguang Chen, Bofei Liu, Jiannan Guan 외 arxiv

Ensuring large language model (LLM) reliability requires distinguishing objective unsolvability (inherent contradictions) from subjective capability limitations (tasks exceeding model competence). Current LLMs often conf…

Reinforcement Learning

Exploring Inevitable Waypoints for Unsolvability Explanation in Hybrid Planning Problems

2025-04-22 · Mir Md Sajid Sarwar, Rajarshi Ray

Explaining unsolvability of planning problems is of significant research interest in Explainable AI Planning. AI planning literature has reported several research efforts on generating explanations of solutions to planni…

Philosophy

Scaling Enterprise Agent Routing: Degradation, Diagnosis, and Recovery

2026-06-16 · Kellen Gillespie, Robyn Perry arxiv

Production LLM assistants route user requests to growing libraries of specialized tools, but how does routing accuracy degrade as the catalog scales? We study single-step routing on a 110-agent, 584-tool catalog from a d…

Mechanism-level routing failure in LLMs over Lean-verified algebraic structures

2026-07-05 · Manuel Israel Cázares, Wenlin Zhang, Haobo Ma arxiv

We present an empirical study of structural routing failure in large language models (LLMs) over a formally verified algebraic corpus. The task requires selecting the correct proof-mechanism label from a fixed closed tem…