paper-with-me

홈 › Papers

CLASSLA-Express: a Train of CLARIN.SI Workshops on Language Resources and Tools with Easily Expanding Route

2024-12-02 · Nikola Ljubešić, Taja Kuzman, Ivana Filipović Petrović, Jelena Parizoska, Petya Osenova

This paper introduces the CLASSLA-Express workshop series as an innovative approach to disseminating linguistic resources and infrastructure provided by the CLASSLA Knowledge Centre for South Slavic languages and the Slovenian CLARIN.SI infrastructure. The workshop series employs two key strategies: (1) conducting workshops directly in countries with interested audiences, and (2) designing the series for easy expansion to new venues. The first iteration of the CLASSLA-Express workshop series encompasses 6 workshops in 5 countries. Its goal is to share knowledge on the use of corpus querying tools, as well as the recently-released CLASSLA-web corpora - the largest general corpora for South Slavic languages. In the paper, we present the design of the workshop series, its current scope and the effortless extensions of the workshop to new venues that are already in sight.

📄 PDF Abstract BibTeX arXiv:2412.01386

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A CLARIN Transcription Portal for Interview Data

2020-05-01 · LREC 2020 5 · Christoph Draxler, Henk van den Heuvel, Arjan van Hessen, Silvia Calamai 외

In this paper we present a first version of a transcription portal for audio files based on automatic speech recognition (ASR) in various languages. The portal is implemented in the CLARIN resources research network and …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

CLASSLA-Stanza: The Next Step for Linguistic Processing of South Slavic Languages

2023-08-08 · Luka Terčon, Nikola Ljubešić

We present CLASSLA-Stanza, a pipeline for automatic linguistic annotation of the South Slavic languages, which is based on the Stanza natural language processing pipeline. We describe the main improvements in CLASSLA-Sta…

The Growing Gains and Pains of Iterative Web Corpora Crawling: Insights from South Slavic CLASSLA-web 2.0 Corpora

2026-01-16 · Taja Kuzman Pungeršek, Peter Rupnik, Vít Suchomel, Nikola Ljubešić arxiv

Crawling national top-level domains has proven to be highly effective for collecting texts in less-resourced languages. This approach has been recently used for South Slavic languages and resulted in the largest general …

Inforex --- a collaborative system for text corpora annotation and analysis

2017-09-01 · RANLP 2017 9 · Micha{\l} Marci{\'n}czuk, Marcin Oleksy, Jan Koco{\'n}

We report a first major upgrade of Inforex {---} a web-based system for qualitative and collaborative text corpora annotation and analysis. Inforex is a part of Polish CLARIN infrastructure. It is integrated with a digit…

Named Entity Recognition (NER)Word Sense Disambiguation

Embedding Software Intent: Lightweight Java Module Recovery

2025-12-17 · Yirui He, Yuqi Huai, Xingyu Chen, Joshua Garcia arxiv

As an increasing number of software systems reach unprecedented scale, relying solely on code-level abstractions is becoming impractical. While architectural abstractions offer a means to manage these systems, maintainin…