paper-with-me

홈 › Papers

Stronger Random Baselines for In-Context Learning

2024-04-19 · Gregory Yauney, David Mimno

Evaluating the in-context learning classification performance of language models poses challenges due to small dataset sizes, extensive prompt-selection using the validation set, and intentionally difficult tasks that lead to near-random performance. The standard random baseline--the expected accuracy of guessing labels uniformly at random--is stable when the evaluation set is used only once or when the dataset is large. We account for the common practice of validation set reuse and existing small datasets with a stronger random baseline: the expected maximum accuracy across multiple random classifiers. When choosing the best prompt demonstrations across six quantized language models applied to 16 BIG-bench Lite tasks, more than 20% of the few-shot results that exceed the standard baseline do not exceed this stronger random baseline. When held-out test sets are available, this stronger baseline is also a better predictor of held-out performance than the standard baseline, avoiding unnecessary test set evaluations. This maximum random baseline provides an easily calculated drop-in replacement for the standard baseline.

📄 PDF Abstract BibTeX arXiv:2404.13020

Code (2)

gyauney/max-random-baseline 공식 구현
gyauney/stronger-random-baselines 공식 구현

Tasks

In-Context Learning

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Stronger Baselines for Trustable Results in Neural Machine Translation

2017-06-29 · WS 2017 8 · Michael Denkowski, Graham Neubig

Interest in neural machine translation has grown rapidly as its effectiveness has been demonstrated across language and data scenarios. New research regularly introduces architectural and algorithmic improvements that le…

Machine TranslationNMTTranslation

InDEX: Indonesian Idiom and Expression Dataset for Cloze Test

2022-11-24 · Xinying Qiu, Guofeng Shi

We propose InDEX, an Indonesian Idiom and Expression dataset for cloze test. The dataset contains 10438 unique sentences for 289 idioms and expressions for which we generate 15 different types of distractors, resulting i…

Cloze TestReading Comprehension

Understanding Context Sampling in TabPFN on Small Tabular Datasets

2026-07-29 · Mohammed Abdullah arxiv

TabPFN performs classification through in-context learning: it conditions on a set of labeled training rows (the context, or prototypes) and predicts test labels without gradient updates. On small tabular datasets, pract…

GANMEX: One-vs-One Attributions Guided by GAN-based Counterfactual Explanation Baselines

2020-11-11 · Sheng-Min Shih, Pin-Ju Tien, Zohar Karnin

Attribution methods have been shown as promising approaches for identifying key features that led to learned model predictions. While most existing attribution methods rely on a baseline input for performing feature pert…

counterfactualCounterfactual Explanation

GANMEX: Class-Targeted One-vs-One Attributions using GAN-based Model Explainability

2021-01-01 · Sheng-Min Shih, Pin-Ju Tien, Zohar Karnin

Attribution methods have been shown as promising approaches for identifying key features that led to learned model predictions. While most existing attribution methods rely on a baseline input for performing feature pert…