paper-with-me

홈 › Papers

Modeling a Historical Variety of a Low-Resource Language: Language Contact Effects in the Verbal Cluster of Early-Modern Frisian

2019-08-01 · WS 2019 8 · Jelke Bloem, Arjen Versloot, Fred Weerman

Certain phenomena of interest to linguists mainly occur in low-resource languages, such as contact-induced language change. We show that it is possible to study contact-induced language change computationally in a historical variety of a low-resource language, Early-Modern Frisian, by creating a model using features that were established to be relevant in a closely related language, modern Dutch. This allows us to test two hypotheses on two types of language contact that may have taken place between Frisian and Dutch during this time. Our model shows that Frisian verb cluster word orders are associated with different context features than Dutch verb orders, supporting the {`}learned borrowing{'} hypothesis.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Seventeenth-Century Spanish American Notary Records for Fine-Tuning Spanish Large Language Models

2024-06-09 · Shraboni Sarker, Ahmad Tamim Hamad, Hulayyil Alshammari, Viviana Grieco 외

Large language models have gained tremendous popularity in domains such as e-commerce, finance, healthcare, and education. Fine-tuning is a common approach to customize an LLM on a domain-specific dataset for a desired d…

Language ModelingLanguage ModellingMasked Language Modeling

Learning from Within? Comparing PoS Tagging Approaches for Historical Text

2016-05-01 · LREC 2016 5 · Sarah Schulz, Jonas Kuhn

In this paper, we investigate unsupervised and semi-supervised methods for part-of-speech (PoS) tagging in the context of historical German text. We locate our research in the context of Digital Humanities where the non-…

Part-Of-Speech TaggingPOSPOS TaggingSelf-Learning

Language Resources for Historical Newspapers: the Impresso Collection

2020-05-01 · LREC 2020 5 · Maud Ehrmann, Matteo Romanello, Simon Clematide, Phillip Benjamin Str{\"o}bel 외

Following decades of massive digitization, an unprecedented amount of historical document facsimiles can now be retrieved and accessed via cultural heritage online portals. If this represents a huge step forward in terms…

InteChar: A Unified Oracle Bone Character List for Ancient Chinese Language Modeling

2025-08-12 · Xiaolei Diao, Zhihan Zhou, Lida Shi, Ting Wang 외 arxiv

Constructing historical language models (LMs) plays a crucial role in aiding archaeological provenance studies and understanding ancient cultures. However, existing resources present major challenges for training effecti…

Data Augmentation

When Alignment Hurts: Decoupling Representational Spaces in Multilingual Models

2025-08-18 · Ahmed Elshabrawy, Hour Kaing, Haiyue Song, Alham Fikri Aji 외 arxiv

Alignment with high-resource standard languages is often assumed to aid the modeling of related low-resource varieties. We challenge this assumption by demonstrating that excessive representational entanglement with a do…