paper-with-me

홈 › Papers

Lexicalized Constituency Parsing for Middle Dutch: Low-resource Training and Cross-Domain Generalization

2026-01-11 · Yiming Liang, Fang Zhao arxiv

Recent years have seen growing interest in applying neural networks and contextualized word embeddings to the parsing of historical languages. However, most advances have focused on dependency parsing, while constituency parsing for low-resource historical languages like Middle Dutch has received little attention. In this paper, we adapt a transformer-based constituency parser to Middle Dutch, a highly heterogeneous and low-resource language, and investigate methods to improve both its in-domain and cross-domain performance. We show that joint training with higher-resource auxiliary languages increases F1 scores by up to 0.73, with the greatest gains achieved from languages that are geographically and temporally closer to Middle Dutch. We further evaluate strategies for leveraging newly annotated data from additional domains, finding that fine-tuning and data combination yield comparable improvements, and our neural parser consistently outperforms the currently used PCFG-based parser for Middle Dutch. We further explore feature-separation techniques for domain adaptation and demonstrate that a minimum threshold of approximately 200 examples per domain is needed to effectively enhance cross-domain performance.

📄 PDF Abstract BibTeX arXiv:2601.07008

Code (0)

등록된 구현이 없습니다.

Tasks

Domain GeneralizationConstituency ParsingDependency ParsingDomain Adaptation

Similar Papers 제목 키워드 기반

Cross-Lingual Constituency Parsing for Middle High German: A Delexicalized Approach

2023-08-09 · Ercong Nie, Helmut Schmid, Hinrich Schütze

Constituency parsing plays a fundamental role in advancing natural language processing (NLP) tasks. However, training an automatic syntactic analysis system for ancient languages solely relying on annotated parse data is…

Constituency ParsingCross-Lingual Transfer

Unlexicalized Transition-based Discontinuous Constituency Parsing

2019-02-24 · TACL 2019 3 · Maximin Coavoux, Benoît Crabbé, Shay B. Cohen

Lexicalized parsing models are based on the assumptions that (i) constituents are organized around a lexical head (ii) bilexical statistics are crucial to solve ambiguities. In this paper, we introduce an unlexicalized t…

Constituency Parsing

Multilingual Lexicalized Constituency Parsing with Word-Level Auxiliary Tasks

2017-04-01 · EACL 2017 4 · Maximin Coavoux, Beno{\^\i}t Crabb{\'e}

We introduce a constituency parser based on a bi-LSTM encoder adapted from recent work (Cross and Huang, 2016b; Kiperwasser and Goldberg, 2016), which can incorporate a lower level character biLSTM (Ballesteros et al., 2…

Constituency ParsingMorphological AnalysisMorphological TaggingPOS

Nested Named Entity Recognition as Latent Lexicalized Constituency Parsing

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Nested named entity recognition (NER) has been receiving increasing attention. Recently, Fu et al. (2020) adapt a span-based constituency parser to tackle nested NER. They treat nested entities as partially-observed cons…

Constituency ParsingEntity Typingnamed-entity-recognitionNamed Entity Recognition+3

Nested Named Entity Recognition as Latent Lexicalized Constituency Parsing

2022-03-09 · ACL 2022 5 · Chao Lou, Songlin Yang, Kewei Tu

Nested named entity recognition (NER) has been receiving increasing attention. Recently, (Fu et al, 2021) adapt a span-based constituency parser to tackle nested NER. They treat nested entities as partially-observed cons…

Constituency ParsingEntity Typingnamed-entity-recognitionNamed Entity Recognition+3