paper-with-me

Papers

When silver glitters more than gold: Bootstrapping an Italian part-of-speech tagger for Twitter

2016-11-09 · Barbara Plank, Malvina Nissim

We bootstrap a state-of-the-art part-of-speech tagger to tag Italian Twitter data, in the context of the Evalita 2016 PoSTWITA shared task. We show that training the tagger on native Twitter data enriched with little amounts of specifically selected gold data and additional silver-labelled data scraped from Facebook, yields better results than using large amounts of manually annotated data from a mix of genres.

📄 PDF Abstract BibTeX arXiv:1611.03057

Code (0)

등록된 구현이 없습니다.

Tasks

TAG

Similar Papers 제목 키워드 기반

Turning silver into gold: error-focused corpus reannotation with active learning

2019-09-01 · RANLP 2019 9 · Pierre Andr{\'e} M{\'e}nard, Antoine Mougeot

While high quality gold standard annotated corpora are crucial for most tasks in natural language processing, many annotated corpora published in recent years, created by annotators or tools, contains noisy annotations. …

Active LearningDocument ClassificationPart-Of-Speech Tagging

SilverAlign: MT-Based Silver Data Algorithm For Evaluating Word Alignment

2022-10-12 · Abdullatif Köksal, Silvia Severini, Hinrich Schütze

Word alignments are essential for a variety of NLP tasks. Therefore, choosing the best approaches for their creation is crucial. However, the scarce availability of gold evaluation data makes the choice difficult. We pro…

Machine TranslationTranslationvalidWord Alignment

False Confidence: Automated Labels Confound Fairness Audits in Cervical Spine Segmentation

2026-07-08 · Linus Juni, Aasa Feragen, Aditya Parikh arxiv

Automated segmentation of cervical-spine MRI is increasingly used in clinical workflows, yet no fairness audit exists for this anatomy. We show that auditing these segmentation tasks is complicated by a common property o…

Cross-Lingual Transfer for Distantly Supervised and Low-resources Indonesian NER

2019-07-25 · Fariz Ikhwantri

Manually annotated corpora for low-resource languages are usually small in quantity (gold), or large but distantly supervised (silver). Inspired by recent progress of injecting pre-trained language model (LM) on many Nat…

Cross-Lingual TransferLanguage ModelingLanguage ModellingNER+2

Centroids: Gold standards with distributional variation

2012-05-01 · LREC 2012 5 · Ian Lewin, {\c{S}}enay Kafkas, Dietrich Rebholz-Schuhmann

Motivation: Gold Standards for named entities are, ironically, not standard themselves. Some specify the “one perfect annotation”. Others specify “perfectly good alternatives”. The concept of Silver standard is relativel…

Named Entity Recognition (NER)