paper-with-me

홈 › Papers

Training Subset Selection for Weak Supervision

2022-06-06 · Hunter Lang, Aravindan Vijayaraghavan, David Sontag

Existing weak supervision approaches use all the data covered by weak signals to train a classifier. We show both theoretically and empirically that this is not always optimal. Intuitively, there is a tradeoff between the amount of weakly-labeled data and the precision of the weak labels. We explore this tradeoff by combining pretrained data representations with the cut statistic (Muhlenbach et al., 2004) to select (hopefully) high-quality subsets of the weakly-labeled training data. Subset selection applies to any label model and classifier and is very simple to plug in to existing weak supervision pipelines, requiring just a few lines of code. We show our subset selection method improves the performance of weak supervision for a wide range of label models, classifiers, and datasets. Using less weakly-labeled data improves the accuracy of weak supervision pipelines by up to 19% (absolute) on benchmark tasks.

📄 PDF Abstract BibTeX arXiv:2206.02914

Code (1)

hunterlang/weaksup-subset-selection 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Optimizing What We Trust: Reliability-Guided QUBO Selection of Multi-Agent Weak Framing Signals for Arabic Sentiment Prediction

2026-02-04 · Rabab Alkhalifa arxiv

Framing detection in Arabic social media is difficult due to interpretive ambiguity, cultural grounding, and limited reliable supervision. Existing LLM-based weak supervision methods typically rely on label aggregation, …

Meta-Learning for Neural Relation Classification with Distant Supervision

2020-10-26 · Zhenzhen Li, Jian-Yun Nie, Benyou Wang, Pan Du 외

Distant supervision provides a means to create a large number of weakly labeled data at low cost for relation classification. However, the resulting labeled instances are very noisy, containing data with wrong labels. Ma…

ClassificationGeneral ClassificationMeta-LearningRelation+1

Semi-Supervised Data Programming with Subset Selection

2020-08-22 · Findings (ACL) 2021 8 · Ayush Maheshwari, Oishik Chatterjee, KrishnaTeja Killamsetty, Ganesh Ramakrishnan 외

The paradigm of data programming, which uses weak supervision in the form of rules/labelling functions, and semi-supervised learning, which augments small amounts of labelled data with a large unlabelled dataset, have sh…

text-classificationText Classification

Socratic Learning: Augmenting Generative Models to Incorporate Latent Subsets in Training Data

2016-10-25 · Paroma Varma, Bryan He, Dan Iter, Peng Xu 외

A challenge in training discriminative models like neural networks is obtaining enough labeled training data. Recent approaches use generative models to combine weak supervision sources, like user-defined heuristics or k…

Relation Extraction

Improving Large-Scale Weakly Supervised ASR by Filtering and Selection

2026-06-27 · Kohei Matsuura, Masato Mimura arxiv

Leveraging large-scale weakly supervised datasets is crucial to train robust end-to-end automatic speech recognition (ASR) models. However, such datasets often contain noisy labels and lack domain specificity, limiting t…

Speech Recognition