paper-with-me

홈 › Papers

EverydayGPT: Confidence-Gated Routing for Efficient and Safe Hybrid GPT-RAG Conversational QA

2026-04-24 · Jaspreet Singh Nahal arxiv

Standard Retrieval-Augmented Generation (RAG) pipelines route every query through retrieval and generation unconditionally, incurring unnecessary computation and propagating low-quality context to the generator. We introduce EverydayGPT, a lightweight conversational QA system built around a Confidence-Gated Routing (CGR) mechanism that formalises the routing decision as a joint policy over retrieval distance and extraction adequacy. The backbone is a 205M-parameter GPT trained from scratch on 10B tokens of FineWeb-Edu. CGR avoids invoking the costly GPT pathway (~5.9s) for 85 percent of queries by resolving them via fast RAG extraction (~45 ms), yielding over 120x latency reduction on the majority of queries while maintaining answer quality. On a 500-question in-domain benchmark, the system achieves F1 = 0.226 +/- 0.004 compared to 0.171 for GPT-only and 0.210 for unconditional RAG. Gains over strong baselines are modest but consistent, while efficiency improvements are substantial (6.3x mean latency reduction). A structured grounding audit finds no unsupported claims in the sampled set, with explicit scope limitations. We position this work as a study of routing strategies under resource constraints rather than a claim of state-of-the-art performance.

📄 PDF Abstract BibTeX arXiv:2606.11212

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ThinkRouter: Efficient Reasoning via Routing Thinking between Latent and Discrete Spaces

2026-02-12 · Xin Xu, Tong Yu, Xiang Chen, Haoliang Wang 외 arxiv

Recent work explores latent reasoning to improve reasoning efficiency by replacing explicit reasoning trajectories with continuous representations in a latent space, yet its effectiveness varies across settings. Analysis…

Mixture of Layers with Hybrid Attention

2026-05-10 · Ivan Ternovtsii, Yurii Bilak arxiv

Standard Mixture-of-Experts (MoE) transformers route tokens to expert subnetworks within each layer, but the layer structure itself remains monolithic. We introduce Mixture of Layers (MoL), which replaces full-width tran…

Resilient Routing: Risk-Aware Dynamic Routing in Smart Logistics via Spatiotemporal Graph Learning

2026-01-20 · Zhiming Xue, Sichen Zhao, Yalun Qi, Xianling Zeng 외 arxiv

With the rapid development of the e-commerce industry, the logistics network is experiencing unprecedented pressure. The traditional static routing strategy most time cannot tolerate the traffic congestion and fluctuatin…

Graph Learning

Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing

2026-07-08 · Tommaso Cerruti, Tim Rieder, George Rowlands, Lingfeng Jin 외 arxiv

Self-attention lets each token retrieve information from the full context, but its quadratic cost in sequence length limits training and inference at long context. This paper presents a comparative study of softmax atten…

When to Think Fast and Slow? AMOR: Adaptive Entropy Gate for Hybrid Models

2026-01-22 · Haoran Zheng, Chen Shani arxiv

Recurrent-attention hybrids aim to combine the efficiency of recurrence with the expressivity of attention, but existing approaches typically apply attention uniformly across all positions, even when the recurrent state …