paper-with-me

홈 › Papers

Icelandic Parallel Abstracts Corpus

2021-08-11 · Haukur Barri Símonarson, Vésteinn Snæbjarnarson

We present a new Icelandic-English parallel corpus, the Icelandic Parallel Abstracts Corpus (IPAC), composed of abstracts from student theses and dissertations. The texts were collected from the Skemman repository which keeps records of all theses, dissertations and final projects from students at Icelandic universities. The corpus was aligned based on sentence-level BLEU scores, in both translation directions, from NMT models using Bleualign. The result is a corpus of 64k sentence pairs from over 6 thousand parallel abstracts.

📄 PDF Abstract BibTeX arXiv:2108.05289

Code (0)

등록된 구현이 없습니다.

Tasks

NMTSentenceTranslation

Similar Papers 제목 키워드 기반

Compiling and Filtering ParIce: An English-Icelandic Parallel Corpus

2019-09-01 · WS (NoDaLiDa) 2019 9 · Starkaður Barkarson, Steinþór Steingrímsson

We present ParIce, a new English-Icelandic parallel corpus. This is the first parallel corpus built for the purposes of language technology development and research for Icelandic, although some Icelandic texts can be fou…

Creating a Parallel Icelandic Dependency Treebank from Raw Text to Universal Dependencies

2020-05-01 · LREC 2020 5 · Hildur J{\'o}nsd{\'o}ttir, Anton Karl Ingason

Making the low-resource language, Icelandic, accessible and usable in Language Technology is a work in progress and is supported by the Icelandic government. Creating resources and suitable training data (e.g., a depende…

A 500 Million Word POS-Tagged Icelandic Corpus

2014-05-01 · LREC 2014 5 · Thomas Eckart, Erla Hallsteinsd{\'o}ttir, Sigr{\'u}n Helgad{\'o}ttir, Uwe Quasthoff 외

The new POS-tagged Icelandic corpus of the Leipzig Corpora Collection is an extensive resource for the analysis of the Icelandic language. As it contains a large share of all Web documents hosted under the .is top-level …

Part-Of-Speech TaggingPOSTAG

Samrómur Children: An Icelandic Speech Corpus

2022-06-01 · LREC 2022 6 · Carlos Daniel Hernandez Mena, David Erik Mollberg, Michal Borský, Jón Guðnason

Samrómur Children is an Icelandic speech corpus intended for the field of automatic speech recognition. It contains 131 hours of read speech from Icelandic children aged between 4 to 17 years. The test portion was meticu…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

A Warm Start and a Clean Crawled Corpus -- A Recipe for Good Language Models

2022-01-14 · Vésteinn Snæbjarnarson, Haukur Barri Símonarson, Pétur Orri Ragnarsson, Svanhvít Lilja Ingólfsdóttir 외

We train several language models for Icelandic, including IceBERT, that achieve state-of-the-art performance in a variety of downstream tasks, including part-of-speech tagging, named entity recognition, grammatical error…

Constituency ParsingGrammatical Error Detectionnamed-entity-recognitionNamed Entity Recognition+3