paper-with-me

홈 › Papers

HFST-SweNER --- A New NER Resource for Swedish

2014-05-01 · LREC 2014 5 · Dimitrios Kokkinakis, Jyrki Niemi, Sam Hardwick, Krister Lind{\'e}n, Lars Borin

Named entity recognition (NER) is a knowledge-intensive information extraction task that is used for recognizing textual mentions of entities that belong to a predefined set of categories, such as locations, organizations and time expressions. NER is a challenging, difficult, yet essential preprocessing technology for many natural language processing applications, and particularly crucial for language understanding. NER has been actively explored in academia and in industry especially during the last years due to the advent of social media data. This paper describes the conversion, modeling and adaptation of a Swedish NER system from a hybrid environment, with integrated functionality from various processing components, to the Helsinki Finite-State Transducer Technology (HFST) platform. This new HFST-based NER (HFST-SweNER) is a full-fledged open source implementation that supports a variety of generic named entity types and consists of multiple, reusable resource layers, e.g., various n-gram-based named entity lists (gazetteers).

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine Translationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NERQuestion AnsweringRelation ExtractionText Classification

Similar Papers 제목 키워드 기반

Hidden Resources ― Strategies to Acquire and Exploit Potential Spoken Language Resources in National Archives

2016-05-01 · LREC 2016 5 · Jens Edlund, Joakim Gustafson

In 2014, the Swedish government tasked a Swedish agency, The Swedish Post and Telecom Authority (PTS), with investigating how to best create and populate an infrastructure for spoken language resources (Ref N2014/2840/IT…

Position

Extracting Semantic Frames using hfst-pmatch

2015-05-01 · WS 2015 5 · Sam Hardwick, Miikka Silfverberg, Krister Lind{\'e}n

The open lexical infrastructure of Spr\aakbanken

2012-05-01 · LREC 2012 5 · Lars Borin, Markus Forsberg, Leif-J{\"o}ran Olsson, Jonatan Uppstr{\"o}m

We present our ongoing work on Karp, Spr{\aa}kbanken's (the Swedish Language Bank) open lexical infrastructure, which has two main functions: (1) to support the work on creating, curating, and integrating our various lex…

A Digital Swedish-Yiddish/Yiddish-Swedish Dictionary: A Web-Based Dictionary that is also Available Offline

2022-06-01 · EURALI (LREC) 2022 6 · Magnus Ahltorp, Jean Hessel, Gunnar Eriksson, Maria Skeppstedt 외

Yiddish is one of the national minority languages of Sweden, and one of the languages for which the Swedish Institute for Language and Folklore is responsible for developing useful language resources. We here describe th…

Transliteration

An Extensible Multilingual Open Source Lemmatizer

2017-09-01 · RANLP 2017 9 · Ahmet Aker, Johann Petrak, Firas Sabbah

We present GATE DictLemmatizer, a multilingual open source lemmatizer for the GATE NLP framework that currently supports English, German, Italian, French, Dutch, and Spanish, and is easily extensible to other languages. …

Information RetrievalLEMMALemmatization