paper-with-me

홈 › Papers

EstNLTK 1.6: Remastered Estonian NLP Pipeline

2020-05-01 · LREC 2020 5 · Sven Laur, Siim Orasmaa, Dage S{\"a}rg, Paul Tammo

The goal of the EstNLTK Python library is to provide a unified programming interface for natural language processing in Estonian. As such, previous versions of the library have been immensely successful both in academic and industrial circles. However, they also contained serious structural limitations {--} it was hard to add new components and there was a lack of fine-grained control needed for back-end programming. These issues have been explicitly addressed in the EstNLTK library while preserving the intuitive interface for novices. We have remastered the basic NLP pipeline by adding many data cleaning steps that are necessary for analyzing real-life texts, and state of the art components for morphological analysis and fact extraction. Our evaluation on unlabelled data shows that the remastered basic NLP pipeline outperforms both the previous version of the toolkit, as well as neural models of StanfordNLP. In addition, EstNLTK contains a new interface for storing, processing and querying text objects in Postgres database which greatly simplifies processing of large text collections. EstNLTK is freely available under the GNU GPL version 2 license, which is standard for academic software.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Morphological Analysis

Similar Papers 제목 키워드 기반

EstNLTK - NLP Toolkit for Estonian

2016-05-01 · LREC 2016 5 · Siim Orasmaa, Timo Petmanson, Alex Tkachenko, er 외

Although there are many tools for natural language processing tasks in Estonian, these tools are very loosely interoperable, and it is not easy to build practical applications on top of them. In this paper, we introduce …

Morphological Analysisnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)

Evaluating Sentence Segmentation and Word Tokenization Systems on Estonian Web Texts

2020-11-16 · Kairit Sirts, Kairit Peekman

Texts obtained from web are noisy and do not necessarily follow the orthographic sentence and word boundary rules. Thus, sentence segmentation and word tokenization systems that have been developed on well-formed texts m…

SegmentationSentenceSentence segmentation

Estonian Native Large Language Model Benchmark

2025-10-24 · Helena Grete Lillepalu, Tanel Alumäe arxiv

The availability of LLM benchmarks for the Estonian language is limited, and a comprehensive evaluation comparing the performance of different LLMs on Estonian tasks has yet to be conducted. We introduce a new benchmark …

Machine Translation

Improving Estonian Text Simplification through Pretrained Language Models and Custom Datasets

2025-01-26 · Eduard Barbu, Meeri-Ly Muru, Sten Marcus Malva

This study introduces an approach to Estonian text simplification using two model architectures: a neural machine translation model and a fine-tuned large language model (LLaMA). Given the limited resources for Estonian,…

Language ModelingLanguage ModellingLarge Language ModelMachine Translation+2

Interesting cross-border news discovery using cross-lingual article linking and document similarity

2021-04-01 · EACL (Hackashop) 2021 4 · Boshko Koloski, Elaine Zosa, Timen Stepišnik-Perdih, Blaž Škrlj 외

Team Name: team-8 Embeddia Tool: Cross-Lingual Document Retrieval Zosa et al. Dataset: Estonian and Latvian news datasets abstract: Contemporary news media face increasing amounts of available data that can be of use whe…

ArticlesRetrieval