paper-with-me

홈 › Papers

End-to-End Weak Supervision

2021-07-05 · NeurIPS 2021 12 · Salva Rühling Cachay, Benedikt Boecking, Artur Dubrawski

Aggregating multiple sources of weak supervision (WS) can ease the data-labeling bottleneck prevalent in many machine learning applications, by replacing the tedious manual collection of ground truth labels. Current state of the art approaches that do not use any labeled training data, however, require two separate modeling steps: Learning a probabilistic latent variable model based on the WS sources -- making assumptions that rarely hold in practice -- followed by downstream model training. Importantly, the first step of modeling does not consider the performance of the downstream model. To address these caveats we propose an end-to-end approach for directly learning the downstream model by maximizing its agreement with probabilistic labels generated by reparameterizing previous probabilistic posteriors with a neural network. Our results show improved performance over prior work in terms of end model performance on downstream test sets, as well as in terms of improved robustness to dependencies among weak supervision sources.

📄 PDF Abstract BibTeX arXiv:2107.02233

Code (1)

autonlab/weasel 공식 구현 pytorch

Tasks

Classification

Similar Papers 제목 키워드 기반

Detecting Fake News with Weak Social Supervision

2019-10-24 · Kai Shu, Ahmed Hassan Awadallah, Susan Dumais, Huan Liu

Limited labeled data is becoming the largest bottleneck for supervised learning systems. This is especially the case for many real-world tasks where large scale annotated examples are either too expensive to acquire or u…

Fake News Detection

WALNUT: A Benchmark on Semi-weakly Supervised Learning for Natural Language Understanding

2021-08-28 · NAACL 2022 7 · Guoqing Zheng, Giannis Karamanolakis, Kai Shu, Ahmed Hassan Awadallah

Building machine learning models for natural language understanding (NLU) tasks relies heavily on labeled data. Weak supervision has been proven valuable when large amount of labeled data is unavailable or expensive to o…

Natural Language UnderstandingWeakly-supervised Learning

Training Subset Selection for Weak Supervision

2022-06-06 · Hunter Lang, Aravindan Vijayaraghavan, David Sontag

Existing weak supervision approaches use all the data covered by weak signals to train a classifier. We show both theoretically and empirically that this is not always optimal. Intuitively, there is a tradeoff between th…

Passage Ranking with Weak Supervision

2019-05-15 · ICLR Workshop LLD 2019 · Peng Xu, Xiaofei Ma, Ramesh Nallapati, Bing Xiang

In this paper, we propose a \textit{weak supervision} framework for neural ranking tasks based on the data programming paradigm \citep{Ratner2016}, which enables us to leverage multiple weak supervision signals from diff…

Passage Ranking

Bandit Label Inference for Weakly Supervised Learning

2015-09-22 · Ke Li, Jitendra Malik

The scarcity of data annotated at the desired level of granularity is a recurring issue in many applications. Significant amounts of effort have been devoted to developing weakly supervised methods tailored to each indiv…

Weakly-supervised Learning