paper-with-me

Papers

The open lexical infrastructure of Spr\aakbanken

2012-05-01 · LREC 2012 5 · Lars Borin, Markus Forsberg, Leif-J{\"o}ran Olsson, Jonatan Uppstr{\"o}m

We present our ongoing work on Karp, Spr{\aa}kbanken's (the Swedish Language Bank) open lexical infrastructure, which has two main functions: (1) to support the work on creating, curating, and integrating our various lexical resources; and (2) to publish daily versions of the resources, making them searchable and downloadable. An important requirement on the lexical infrastructure is also that we maintain a strong bidirectional connection to our corpus infrastructure. At the heart of the infrastructure is the SweFN++ project with the goal to create free Swedish lexical resources geared towards language technology applications. The infrastructure currently hosts 15 Swedish lexical resources, including historical ones, some of which have been created from scratch using existing free resources, both external and in-house. The resources are integrated through links to a pivot lexical resource, SALDO, a large morphological and lexical-semantic resource for modern Swedish. SALDO has been selected as the pivot partly because of its size and quality, but also because its form and sense units have been assigned persistent identifiers (PIDs) to which the lexical information in other lexical resources and in corpora are linked.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Korp and Karp -- A Bestiary of Language Resources: The Research Infrastructure of Spr\aakbanken

2013-05-01 · WS 2013 5 · Malin Ahlberg, Lars Borin, Markus Forsberg, Martin Hammarstedt 외

Korp --- the corpus infrastructure of Spr\aakbanken

2012-05-01 · LREC 2012 5 · Lars Borin, Markus Forsberg, Johan Roxendal

We present Korp, the corpus infrastructure of Spr{\aa}kbanken (the Swedish Language Bank). The infrastructure consists of three main components: the Korp corpus pipeline, the Korp backend, and the Korp frontend. The Korp…

Lemmatization

Open-Source Morphology for Endangered Mordvinic Languages

2020-11-11 · Jack Rueter, Mika Hämäläinen, Niko Partanen

This document describes shared development of finite-state description of two closely related but endangered minority languages, Erzya and Moksha. It touches upon morpholexical unity and diversity of the two languages an…

DiversityUnity

Analyzing and Encoding the Al-Mawrid Arabic-English Dictionary with the ISO Language Markup Framework and TEI Lex-0

2026-06-16 · Diaa M. Fayed, Laurent Romary arxiv

This paper presents a robust methodology for the systematic digitization and encoding of the Al-Mawrid Arabic-English dictionary, transforming it from a legacy print resource into a standardized computational lexicon. Ad…

Information Extraction

A New Integrated Open-source Morphological Analyzer for Hungarian

2016-05-01 · LREC 2016 5 · Attila Nov{\'a}k, Borb{\'a}la Sikl{\'o}si, Csaba Oravecz

The goal of a Hungarian research project has been to create an integrated Hungarian natural language processing framework. This infrastructure includes tools for analyzing Hungarian texts, integrated into a standardized …