paper-with-me

홈 › Papers

Sim-to-Real Betting on the E-Process: Bringing "simulators" to anytime-valid confidence sequences

2026-06-23 · Yujia Chen, Bowen Weng arxiv

This note describes an integration of the sim-to-real performance estimate with betting (from Chen et al.) and the safe anytime-valid inference (from Ramdas et al.). Using the scaled simulators. The method produces efficient, reliable certificates for the mean estimate, an approach that is especially valuable in robot performance testing. This note gives a primary, self-contained account of the construction; preliminaries of the respective methods are kept at a minimum, and one shall refer to the original works for full detail. Some synthetic examples demonstrating the proposed algorithm can be found at https://github.com/ISUSAIL/Bet4Sim2Real-EProcess.

📄 PDF Abstract BibTeX arXiv:2606.24038

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Learning to Bet for Horizon-Aware Anytime-Valid Testing

2026-03-20 · Ege Onur Taga, Samet Oymak, Shubhanshu Shekhar arxiv

We develop horizon-aware anytime-valid tests and confidence sequences for bounded means under a strict deadline $N$. Using the betting/e-process framework, we cast horizon-aware betting as a finite-horizon optimal contro…

Reinforcement Learning

Betting for Sim-to-Real Performance Evaluation

2026-04-27 · Zaid Mahboob, Yujia Chen, Bowen Weng arxiv

This paper studies the problem of robot performance evaluation, focusing on how to obtain accurate and efficient estimates of real-world behavior under severe constraints on physical experimentation. Such estimates are e…

Anytime-Valid Federated Conformal RAG for LLM Swarms

2026-05-27 · Prasanjit Dubey, Xiaoming Huo arxiv

Federated Conformal RAG (FC-RAG) provides distribution-free coverage for a bandwidth-limited swarm of weak language models, but only at a fixed horizon. We extend it to anytime-valid sequential coverage: validity at ever…

Evidence Before Expansion: Reuse, Spawn, or Defer in Lifelong Expert Pools

2026-08-20 · Kentaro Oda arxiv

Streaming systems that maintain a pool of expert models must repeatedly decide whether to reuse an existing expert for arriving data, spawn a new one, or defer. We present a decision layer that makes all three outcomes s…

Auditing Fairness by Betting

2023-05-27 · NeurIPS 2023 11 · Ben Chugg, Santiago Cortes-Gomez, Bryan Wilder, Aaditya Ramdas

We provide practical, efficient, and nonparametric methods for auditing the fairness of deployed classification and regression models. Whereas previous work relies on a fixed-sample size, our methods are sequential and a…

Fairnessvalid