EstNLTK 1.6: Remastered Estonian NLP Pipeline
The goal of the EstNLTK Python library is to provide a unified programming interface for natural language processing in Estonian. As such, previous versions of the library have been immensely successful both in academic and industrial circles. However, they also contained serious structural limitations {--} it was hard to add new components and there was a lack of fine-grained control needed for back-end programming. These issues have been explicitly addressed in the EstNLTK library while preserving the intuitive interface for novices. We have remastered the basic NLP pipeline by adding many data cleaning steps that are necessary for analyzing real-life texts, and state of the art components for morphological analysis and fact extraction. Our evaluation on unlabelled data shows that the remastered basic NLP pipeline outperforms both the previous version of the toolkit, as well as neural models of StanfordNLP. In addition, EstNLTK contains a new interface for storing, processing and querying text objects in Postgres database which greatly simplifies processing of large text collections. EstNLTK is freely available under the GNU GPL version 2 license, which is standard for academic software.
Code (0)
등록된 구현이 없습니다.
Tasks
Morphological AnalysisSimilar Papers 제목 키워드 기반
EstNLTK - NLP Toolkit for Estonian
Although there are many tools for natural language processing tasks in Estonian, these tools are very loosely interoperable, and it is not easy to build practical applications on top of them. In this paper, we introduce …
Morphological Analysisnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Evaluating Sentence Segmentation and Word Tokenization Systems on Estonian Web Texts
Texts obtained from web are noisy and do not necessarily follow the orthographic sentence and word boundary rules. Thus, sentence segmentation and word tokenization systems that have been developed on well-formed texts m…
SegmentationSentenceSentence segmentationEstonian Native Large Language Model Benchmark
The availability of LLM benchmarks for the Estonian language is limited, and a comprehensive evaluation comparing the performance of different LLMs on Estonian tasks has yet to be conducted. We introduce a new benchmark …
Machine TranslationImproving Estonian Text Simplification through Pretrained Language Models and Custom Datasets
This study introduces an approach to Estonian text simplification using two model architectures: a neural machine translation model and a fine-tuned large language model (LLaMA). Given the limited resources for Estonian,…
Language ModelingLanguage ModellingLarge Language ModelMachine Translation+2Interesting cross-border news discovery using cross-lingual article linking and document similarity
Team Name: team-8 Embeddia Tool: Cross-Lingual Document Retrieval Zosa et al. Dataset: Estonian and Latvian news datasets abstract: Contemporary news media face increasing amounts of available data that can be of use whe…
ArticlesRetrieval