paper-with-me

홈 › Papers

PABI: A Unified PAC-Bayesian Informativeness Measure for Incidental Supervision Signals

2021-01-01 · Hangfeng He, Mingyuan Zhang, Qiang Ning, Dan Roth

Real-world applications often require making use of {\em a range of incidental supervision signals}. However, we currently lack a principled way to measure the benefit an incidental training dataset can bring, and the common practice of using indirect, weaker signals is through exhaustive experiments with various models and hyper-parameters. This paper studies whether we can, {\em in a single framework, quantify the benefit of various types of incidental signals for one's target task without going through combinatorial experiments}. We propose PABI, a unified informativeness measure backed by PAC-Bayesian theory, characterizing the reduction in uncertainty that indirect, weak signals provide. We demonstrate PABI's use in quantifying various types of incidental signals including partial labels, noisy labels, constraints, cross-domain signals, and combinations of these. Experiments with various setups on two natural language processing (NLP) tasks, named entity recognition (NER) and question answering (QA), show that PABI correlates well with learning performance, providing a promising way to determine, ahead of learning, which supervision signals would be beneficial.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Informativenessnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NERQuestion Answering

Similar Papers 제목 키워드 기반

Foreseeing the Benefits of Incidental Supervision

2020-06-09 · EMNLP 2021 11 · Hangfeng He, Mingyuan Zhang, Qiang Ning, Dan Roth

Real-world applications often require improved models by leveraging a range of cheap incidental supervision signals. These could include partial labels, noisy labels, knowledge-based constraints, and cross-domain or cros…

InformativenessLearning Theorynamed-entity-recognitionNamed Entity Recognition+3

Laplace Sample Information: Data Informativeness Through a Bayesian Lens

2025-05-21 · Johannes Kaiser, Kristian Schwethelm, Daniel Rueckert, Georgios Kaissis

Accurately estimating the informativeness of individual samples in a dataset is an important objective in deep learning, as it can guide sample selection, which can improve model efficiency and accuracy by removing redun…

Informativeness

Searching for Needles in a Haystack: On the Role of Incidental Bilingualism in PaLM's Translation Capability

2023-05-17 · Eleftheria Briakou, Colin Cherry, George Foster

Large, multilingual language models exhibit surprisingly good zero- or few-shot machine translation capabilities, despite having never seen the intentionally-included translation examples provided to typical neural trans…

Language ModelingLanguage ModellingMachine TranslationTranslation

Proper Dataset Valuation by Pointwise Mutual Information

2024-05-28 · Shuran Zheng, Xuan Qi, Rui Ray Chen, Yongchan Kwon 외

Data plays a central role in advancements in modern artificial intelligence, with high-quality data emerging as a key driver of model performance. This has prompted the development of principled and effective data curati…

Data ValuationInformativeness

When redundancy is useful: A Bayesian approach to 'overinformative' referring expressions

2019-03-19 · Judith Degen, Robert D. Hawkins, Caroline Graf, Elisa Kreiss 외

Referring is one of the most basic and prevalent uses of language. How do speakers choose from the wealth of referring expressions at their disposal? Rational theories of language use have come under attack for decades f…

InformativenessSpecificity