paper-with-me

Papers

Batch Universal Prediction

2024-02-06 · Marco Bondaschi, Michael Gastpar

Large language models (LLMs) have recently gained much popularity due to their surprising ability at generating human-like English sentences. LLMs are essentially predictors, estimating the probability of a sequence of words given the past. Therefore, it is natural to evaluate their performance from a universal prediction perspective. In order to do that fairly, we introduce the notion of batch regret as a modification of the classical average regret, and we study its asymptotical value for add-constant predictors, in the case of memoryless sources and first-order Markov sources.

📄 PDF Abstract BibTeX arXiv:2402.03901

Code (0)

등록된 구현이 없습니다.

Tasks

Prediction

Similar Papers 제목 키워드 기반

Universal distribution of the empirical coverage in split conformal prediction

2023-03-05 · Paulo C. Marques F

When split conformal prediction operates in batch mode with exchangeable data, we determine the exact distribution of the empirical coverage of prediction sets produced for a finite batch of future observables, as well a…

Conformal PredictionPredictionregression

The Conditional Regret-Capacity Theorem for Batch Universal Prediction

2025-08-14 · Marco Bondaschi, Michael Gastpar arxiv

We derive a conditional version of the classical regret-capacity theorem. This result can be used in universal prediction to find lower bounds on the minimal batch regret, which is a recently introduced generalization of…

Universal Supervised Learning for Individual Data

2018-12-22 · Yaniv Fogel, Meir Feder

Universal supervised learning is considered from an information theoretic point of view following the universal prediction approach, see Merhav and Feder (1998). We consider the standard supervised "batch" learning where…

Optimistic Online-to-Batch Conversions for Accelerated Convergence and Universality

2025-11-10 · Yu-Hu Yan, Peng Zhao, Zhi-Hua Zhou arxiv

In this work, we study offline convex optimization with smooth objectives, where the classical Nesterov's Accelerated Gradient (NAG) method achieves the optimal accelerated convergence. Extensive research has aimed to un…

Optimal Learning for Multi-pass Stochastic Gradient Methods

2016-12-01 · NeurIPS 2016 12 · Junhong Lin, Lorenzo Rosasco

We analyze the learning properties of the stochastic gradient method when multiple passes over the data and mini-batches are allowed. In particular, we consider the square loss and show that for a universal step-siz…