paper-with-me

홈 › Papers

Evidence Before Expansion: Reuse, Spawn, or Defer in Lifelong Expert Pools

2026-08-20 · Kentaro Oda arxiv

Streaming systems that maintain a pool of expert models must repeatedly decide whether to reuse an existing expert for arriving data, spawn a new one, or defer. We present a decision layer that makes all three outcomes statistically meaningful. Reuse and spawn are posed as one-sided sequential hypotheses on a conditional (mechanism-level) discrepancy, separated by an indifference zone; defer is exactly the state in which neither betting e-process has accumulated sufficient evidence. We prove finite-time anytime validity for the observable surrogate discrepancy of a predictable discriminator sequence, and an unconditional one-sided transfer to the population quantity in which each side's slack is the excess risk of a single discriminator; an empirically observed downward-bias regularity makes the spawn side exactly conservative. Recency without sacrificing the guarantee is obtained by a restarted e-detector: a bank of unwindowed betting supermartingales at geometrically spaced restart times (O(log t) memory), with the error budget spent over restart instances, which preserves lifetime anytime validity; spending over expert-creation order likewise controls multiplicity for unboundedly many experts. On synthetic multi-concept streams, Electricity, Covertype, and the recurrence-heavy INSECTS benchmark, the instance-accounted restarted bank achieves zero false spawns and zero false reuses after switches and matches or exceeds the retired windowed heuristic (INSECTS-reoccurring accuracy 0.675), making the deployed algorithm and the guaranteed algorithm one and the same.

📄 PDF Abstract BibTeX arXiv:2608.19888

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DeferMem: Query-Time Evidence Distillation via Reinforcement Learning for Long-Term Memory QA

2026-05-21 · Jianing Yin, Tan Tang arxiv

Large language model (LLM) agents still struggle with long-term memory question answering, where answer-supporting evidence is often scattered across long conversational histories and buried in substantial irrelevant con…

Reinforcement LearningQuestion Answering

OmniHarness: Harnessing Generalizable Visual Generation via Symbolic Policy Learning

2026-09-13 · Xu Xu, Jinxiu Liu, Zhangbo Qiao, Jiaxing Lu 외 hf

Unified multimodal large language models (MLLMs) and multi-agent systems have advanced visual generation. However, three limitations remain. (1) Existing methods often distill task-specific experience with limited genera…

Why Ask One When You Can Ask $k$? Two-Stage Learning-to-Defer to the Top-$k$ Experts

2025-04-17 · Yannis Montreuil, Axel Carlier, Lai Xing Ng, Wei Tsang Ooi

Although existing Learning-to-Defer (L2D) frameworks support multiple experts, they allocate each query to a single expert, limiting their ability to leverage collective expertise in complex decision-making scenarios. To…

Decision Making

TripLe: Revisiting Pretrained Model Reuse and Progressive Learning for Efficient Vision Transformer Scaling and Searching

2023-01-01 · ICCV 2023 1 · Cheng Fu, Hanxian Huang, Zixuan Jiang, Yun Ni 외

One promising way to accelerate transformer training is to reuse small pretrained models to initialize the transformer, as their existing representation power facilitates faster model convergence. Previous works desi…

Knowledge DistillationNeural Architecture Search

Don't Commit Alone: Joint Token Commitment in Diffusion Large Language Models

2026-07-05 · Lin Yao arxiv

Diffusion large language models (dLLMs) commit multiple tokens per denoising step by decoding each selected position independently from the shared context; when those positions are dependent, the resulting factorization …