paper-with-me

Papers

Balancing Label Quantity and Quality for Scalable Elicitation

2024-10-17 · Alex Mallen, Nora Belrose

Scalable oversight studies methods of training and evaluating AI systems in domains where human judgment is unreliable or expensive, such as scientific research and software engineering in complex codebases. Most work in this area has focused on methods of improving the quality of labels. Recent work by Burns et al. (2023) considers the complementary problem of training models with low-quality labels, finding that large pretrained models often have an inductive bias towards producing correct answers. In practice, however, neither label quantity nor quality is fixed: practitioners face a quantity-quality tradeoff. In this paper, we explore the microeconomics of the quantity-quality tradeoff on binary NLP classification tasks used in Burns et al. (2023). While sample-efficient learning has been studied extensively, little public research has focused on scalable elicitation: eliciting capabilities from pretrained models subject to labeling cost constraints. We find that this setting has novel dynamics caused by the tradeoff between label quantity and quality, as well as the model's existing latent capabilities. We observe three regimes of eliciting classification knowledge from pretrained models using supervised finetuning: quantity-dominant, quality-dominant, and a mixed regime involving the use of low- and high-quality data together to attain higher accuracy at a lower cost than using either alone. We explore sample-efficient elicitation methods that make use of two datasets of differing qualities, and establish a Pareto frontier of scalable elicitation methods that optimally trade off labeling cost and classifier performance. We find that the accuracy of supervised fine-tuning can be improved by up to 5 percentage points at a fixed labeling budget by adding a few-shot prompt to make use of the model's existing knowledge of the task.

📄 PDF Abstract BibTeX arXiv:2410.13215

Code (1)

eleutherai/scalable-elicitation 공식 구현 pytorch

Tasks

Inductive BiasLanguage Modelling

Similar Papers 제목 키워드 기반

Iterative label cleaning for transductive and semi-supervised few-shot learning

2020-12-14 · ICCV 2021 10 · Michalis Lazarou, Tania Stathaki, Yannis Avrithis

Few-shot learning amounts to learning representations and acquiring knowledge such that novel tasks may be solved with both supervision and data being limited. Improved performance is possible by transductive inference, …

Few-Shot Learning

Scalable Delphi: Large Language Models for Structured Risk Estimation

2026-02-09 · Tobias Lorenz, Mario Fritz arxiv

Quantitative risk assessment in high-stakes domains relies on structured expert elicitation to estimate unobservable properties. The gold standard - the Delphi method - produces calibrated, auditable judgments but requir…

Mixture of Expert/Imitator Networks: Scalable Semi-supervised Learning Framework

2018-10-13 · Shun Kiyono, Jun Suzuki, Kentaro Inui

The current success of deep neural networks (DNNs) in an increasingly broad range of tasks involving artificial intelligence strongly depends on the quality and quantity of labeled training data. In general, the scarcity…

General Classificationtext-classificationText Classification

Acted vs. Improvised: Domain Adaptation for Elicitation Approaches in Audio-Visual Emotion Recognition

2021-04-05 · Haoqi Li, Yelin Kim, Cheng-Hao Kuo, Shrikanth Narayanan

Key challenges in developing generalized automatic emotion recognition systems include scarcity of labeled data and lack of gold-standard references. Even for the cues that are labeled as the same emotion category, the v…

Domain AdaptationEmotion RecognitionTransfer Learning

Quality or Quantity: Toward a Unified Approach for Multi-organ Segmentation in Body CT

2022-03-03 · Fakrul Islam Tushar, Husam Nujaim, Wanyi Fu, Ehsan Abadi 외

Organ segmentation of medical images is a key step in virtual imaging trials. However, organ segmentation datasets are limited in terms of quality (because labels cover only a few organs) and quantity (since case numbers…

Organ SegmentationSegmentation