paper-with-me

홈 › Papers

WikiGoldSK: Annotated Dataset, Baselines and Few-Shot Learning Experiments for Slovak Named Entity Recognition

2023-04-08 · Dávid Šuba, Marek Šuppa, Jozef Kubík, Endre Hamerlik, Martin Takáč

Named Entity Recognition (NER) is a fundamental NLP tasks with a wide range of practical applications. The performance of state-of-the-art NER methods depends on high quality manually anotated datasets which still do not exist for some languages. In this work we aim to remedy this situation in Slovak by introducing WikiGoldSK, the first sizable human labelled Slovak NER dataset. We benchmark it by evaluating state-of-the-art multilingual Pretrained Language Models and comparing it to the existing silver-standard Slovak NER dataset. We also conduct few-shot experiments and show that training on a sliver-standard dataset yields better results. To enable future work that can be based on Slovak NER, we release the dataset, code, as well as the trained models publicly under permissible licensing terms at https://github.com/NaiveNeuron/WikiGoldSK.

📄 PDF Abstract BibTeX arXiv:2304.04026

Code (1)

naiveneuron/wikigoldsk 공식 구현

Tasks

Few-Shot Learningnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER

Similar Papers 제목 키워드 기반

Group-On: Boosting One-Shot Segmentation with Supportive Query

2024-04-18 · Hanjing Zhou, Mingze Yin, Jintai Chen, Danny Chen 외

One-shot semantic segmentation aims to segment query images given only ONE annotated support image of the same class. This task is challenging because target objects in the support and query images can be largely differe…

One-Shot SegmentationSegmentationSemantic Segmentation

Benchmarking zero-shot and few-shot approaches for tokenization, tagging, and dependency parsing of Tagalog text

2022-08-03 · Angelina Aquino, Franz de Leon

The grammatical analysis of texts in any written language typically involves a number of basic processing tasks, such as tokenization, morphological tagging, and dependency parsing. State-of-the-art systems can achieve h…

BenchmarkingData AugmentationDependency ParsingMorphological Tagging+1

ConEntail: An Entailment-based Framework for Universal Zero and Few Shot Classification with Supervised Contrastive Pretraining

2022-10-14 · Ranran Haoran Zhang, Aysa Xuemo Fan, Rui Zhang

A universal classification model aims to generalize to diverse classification tasks in both zero and few shot settings. A promising way toward universal classification is to cast heterogeneous data formats into a dataset…

ClassificationNatural Language InferenceQuestion AnsweringSentence

RelVAE: Generative Pretraining for few-shot Visual Relationship Detection

2023-11-27 · Sotiris Karapiperis, Markos Diomataris, Vassilis Pitsikalis

Visual relations are complex, multimodal concepts that play an important role in the way humans perceive the world. As a result of their complexity, high-quality, diverse and large scale datasets for visual relations are…

Predicate ClassificationRelationship DetectionVisual Relationship Detection

MagicBrush: A Manually Annotated Dataset for Instruction-Guided Image Editing

2023-06-16 · NeurIPS 2023 11 · Kai Zhang, Lingbo Mo, Wenhu Chen, Huan Sun 외

Text-guided image editing is widely needed in daily life, ranging from personal use to professional applications such as Photoshop. However, existing methods are either zero-shot or trained on an automatically synthesize…

Image Editingtext-guided-image-editing