paper-with-me

홈 › Papers

Web Service integration platform for Polish linguistic resources

2012-05-01 · LREC 2012 5 · Maciej Ogrodniczuk, Micha{\l} Lenart

This paper presents a robust linguistic Web service framework for Polish, combining several mature offline linguistic tools in a common online platform. The toolset comprise paragraph-, sentence- and token-level segmenter, morphological analyser, disambiguating tagger, shallow and deep parser, named entity recognizer and coreference resolver. Uniform access to processing results is provided by means of a stand-off packaged adaptation of National Corpus of Polish TEI P5-based representation and interchange format. A concept of asynchronous handling of requests sent to the implemented Web service (Multiservice) is introduced to enable processing large amounts of text by setting up language processing chains of desired complexity. Apart from a dedicated API, a simpleWeb interface to the service is presented, allowing to compose a chain of annotation services, run it and periodically check for execution results, made available as plain XML or in a simple visualization. Usage examples and results from performance and scalability tests are also included.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Sentence

Similar Papers 제목 키워드 기반

Web services and data mining: combining linguistic tools for Polish with an analytical platform

2016-12-01 · WS 2016 12 · Maciej Ogrodniczuk

In this paper we present a new combination of existing language tools for Polish with a popular data mining platform intended to help researchers from digital humanities perform computational analyses without any program…

Integration of Workflow and Pipeline for Language Service Composition

2014-05-01 · LREC 2014 5 · Trang Mai Xuan, Yohei Murakami, Donghui Lin, Toru Ishida

Integrating language resources and language services is a critical part of building natural language processing applications. Service workflow and processing pipeline are two approaches for sharing and combining language…

Service Composition

BAN-PL: a Novel Polish Dataset of Banned Harmful and Offensive Content from Wykop.pl web service

2023-08-21 · Anna Kołos, Inez Okulska, Kinga Głąbińska, Agnieszka Karlińska 외

Since the Internet is flooded with hate, it is one of the main tasks for NLP experts to master automated online content moderation. However, advancements in this field require improved access to publicly available accura…

Specificity

Multisłownik: Linking plWordNet-based Lexical Data for Lexicography and Educational Purposes

2018-01-01 · GWC 2018 1 · Maciej Ogrodniczuk, Joanna Bilińska, Zbigniew Bronk, Witold Kieraś

Multisłownik is an automated integrator of Polish lexical data retrieved from multiple available online sources intended to be used in various scenarios requiring access to such data, most prominently dictionary creation…

Towards a comprehensive open repository of Polish language resources

2012-05-01 · LREC 2012 5 · Maciej Ogrodniczuk, Piotr P{\k{e}}zik, Adam Przepi{\'o}rkowski

The aim of this paper is to present current efforts towards the creation of a comprehensive open repository of Polish language resources and tools (LRTs). The work described here is carried out within the CESAR project, …