paper-with-me

홈 › Papers

A Second Wave of UD Hebrew Treebanking and Cross-Domain Parsing

2022-10-14 · Amir Zeldes, Nick Howell, Noam Ordan, Yifat Ben Moshe

Foundational Hebrew NLP tasks such as segmentation, tagging and parsing, have relied to date on various versions of the Hebrew Treebank (HTB, Sima'an et al. 2001). However, the data in HTB, a single-source newswire corpus, is now over 30 years old, and does not cover many aspects of contemporary Hebrew on the web. This paper presents a new, freely available UD treebank of Hebrew stratified from a range of topics selected from Hebrew Wikipedia. In addition to introducing the corpus and evaluating the quality of its annotations, we deploy automatic validation tools based on grew (Guillaume, 2021), and conduct the first cross domain parsing experiments in Hebrew. We obtain new state-of-the-art (SOTA) results on UD NLP tasks, using a combination of the latest language modelling and some incremental improvements to existing transformer based approaches. We also release a new version of the UD HTB matching annotation scheme updates from our new corpus.

📄 PDF Abstract BibTeX arXiv:2210.07873

Code (2)

amir-zeldes/hebpipe 공식 구현 pytorch
universaldependencies/ud_hebrew-iahltwiki 공식 구현

Tasks

Language Modelling

Similar Papers 제목 키워드 기반

AlephBERT:A Hebrew Large Pre-Trained Language Model to Start-off your Hebrew NLP Application With

2021-04-08 · Amit Seker, Elron Bandel, Dan Bareket, Idan Brusilovsky 외

Large Pre-trained Language Models (PLMs) have become ubiquitous in the development of language understanding technology and lie at the heart of many artificial intelligence advances. While advances reported for English u…

Language ModelingLanguage ModellingMorphological Taggingnamed-entity-recognition+4

Bootstrapping UD treebanks for Delexicalized Parsing

2019-09-01 · WS (NoDaLiDa) 2019 9 · Prasanth Kolachina, Aarne Ranta

Standard approaches to treebanking traditionally employ a waterfall model (Sommerville, 2010), where annotation guidelines guide the annotation process and insights from the annotation process in turn lead to subsequent …

How to construct a multi-lingual domain ontology

2014-05-01 · LREC 2014 5 · Nitsan Chrizman, Alon Itai

The research focuses on automatic construction of multi-lingual domain-ontologies, i.e., creating a DAG (directed acyclic graph) consisting of concepts relating to a specific domain and the relations between them. The do…

ArticlesRelational Reasoning

AlephBERT: Language Model Pre-training and Evaluation from Sub-Word to Sentence Level

2022-05-01 · ACL 2022 5 · Amit Seker, Elron Bandel, Dan Bareket, Idan Brusilovsky 외

Large Pre-trained Language Models (PLMs) have become ubiquitous in the development of language understanding technology and lie at the heart of many artificial intelligence advances. While advances reported for English u…

Language ModelingLanguage ModellingSentence

AlephBERT: Language Model Pre-training and Evaluation from Sub-Word to Sentence Level

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Large Pre-trained Language Models (PLMs) have become ubiquitous in the development of language understanding technology and lie at the heart of many artificial intelligence advances. While advances reported for English u…

Language ModelingLanguage ModellingSentence