paper-with-me

홈 › Papers

CLASSLA-Stanza: The Next Step for Linguistic Processing of South Slavic Languages

2023-08-08 · Luka Terčon, Nikola Ljubešić

We present CLASSLA-Stanza, a pipeline for automatic linguistic annotation of the South Slavic languages, which is based on the Stanza natural language processing pipeline. We describe the main improvements in CLASSLA-Stanza with respect to Stanza, and give a detailed description of the model training process for the latest 2.1 release of the pipeline. We also report performance scores produced by the pipeline for different languages and varieties. CLASSLA-Stanza exhibits consistently high performance across all the supported languages and outperforms or expands its parent pipeline Stanza at all the supported tasks. We also present the pipeline's new functionality enabling efficient processing of web data and the reasons that led to its implementation.

📄 PDF Abstract BibTeX arXiv:2308.04255

Code (1)

clarinsi/classla 공식 구현 pytorch

Similar Papers 제목 키워드 기반

CLASSLA-web: Comparable Web Corpora of South Slavic Languages Enriched with Linguistic and Genre Annotation

2024-03-19 · Nikola Ljubešić, Taja Kuzman

This paper presents a collection of highly comparable web corpora of Slovenian, Croatian, Bosnian, Montenegrin, Serbian, Macedonian, and Bulgarian, covering thereby the whole spectrum of official languages in the South S…

Articles

CLASSLA-Express: a Train of CLARIN.SI Workshops on Language Resources and Tools with Easily Expanding Route

2024-12-02 · Nikola Ljubešić, Taja Kuzman, Ivana Filipović Petrović, Jelena Parizoska 외

This paper introduces the CLASSLA-Express workshop series as an innovative approach to disseminating linguistic resources and infrastructure provided by the CLASSLA Knowledge Centre for South Slavic languages and the Slo…

Stanza: A Python Natural Language Processing Toolkit for Many Human Languages

2020-03-16 · ACL 2020 6 · Peng Qi, Yuhao Zhang, Yuhui Zhang, Jason Bolton 외

We introduce Stanza, an open-source Python natural language processing toolkit supporting 66 human languages. Compared to existing widely used toolkits, Stanza features a language-agnostic fully neural pipeline for text …

Coreference ResolutionDependency ParsingLemmatizationNamed Entity Recognition+2

Extending the SSJ Universal Dependencies Treebank for Slovenian: Was It Worth It?

2022-06-01 · LREC (LAW) 2022 6 · Kaja Dobrovoljc, Nikola Ljubešić

This paper presents the creation and evaluation of a new version of the reference SSJ Universal Dependencies Treebank for Slovenian, which has been substantially improved and extended to almost double the original size. …

Dependency Parsing

The Growing Gains and Pains of Iterative Web Corpora Crawling: Insights from South Slavic CLASSLA-web 2.0 Corpora

2026-01-16 · Taja Kuzman Pungeršek, Peter Rupnik, Vít Suchomel, Nikola Ljubešić arxiv

Crawling national top-level domains has proven to be highly effective for collecting texts in less-resourced languages. This approach has been recently used for South Slavic languages and resulted in the largest general …