paper-with-me

Papers

NoReC: The Norwegian Review Corpus

2017-10-15 · LREC 2018 5 · Erik Velldal, Lilja Øvrelid, Eivind Alexander Bergem, Cathrine Stadsnes, Samia Touileb, Fredrik Jørgensen

This paper presents the Norwegian Review Corpus (NoReC), created for training and evaluating models for document-level sentiment analysis. The full-text reviews have been collected from major Norwegian news sources and cover a range of different domains, including literature, movies, video games, restaurants, music and theater, in addition to product reviews across a range of categories. Each review is labeled with a manually assigned score of 1-6, as provided by the rating of the original author. This first release of the corpus comprises more than 35,000 reviews. It is distributed using the CoNLL-U format, pre-processed using UDPipe, along with a rich set of metadata. The work reported in this paper forms part of the SANT initiative (Sentiment Analysis for Norwegian Text), a project seeking to provide resources and tools for sentiment analysis and opinion mining for Norwegian. As resources for sentiment analysis have so far been unavailable for Norwegian, NoReC represents a highly valuable and sought-after addition to Norwegian language technology.

📄 PDF Abstract BibTeX arXiv:1710.05370

Code (1)

ltgoslo/norec 공식 구현

Tasks

Opinion MiningSentiment Analysis

Similar Papers 제목 키워드 기반

Negation in Norwegian: an annotated dataset

2021-05-01 · NoDaLiDa 2021 5 · Petter Mæhlum, Jeremy Barnes, Robin Kurtz, Lilja Øvrelid 외

This paper introduces NorecNeg – the first annotated dataset of negation for Norwegian. Negation cues and their in-sentence scopes have been annotated across more than 11K sentences spanning more than 400 documents for a…

NegationSentence

A Fine-Grained Sentiment Dataset for Norwegian

2019-11-28 · LREC 2020 5 · Lilja Øvrelid, Petter Mæhlum, Jeremy Barnes, Erik Velldal

We introduce NoReC_fine, a dataset for fine-grained sentiment analysis in Norwegian, annotated with respect to polar expressions, targets and holders of opinion. The underlying texts are taken from a corpus of profession…

Sentiment Analysis

The Norwegian Colossal Corpus: A Text Corpus for Training Large Norwegian Language Models

2022-06-01 · LREC 2022 6 · Per Kummervold, Freddy Wetjen, Javier de la Rosa

Norwegian has been one of many languages lacking sufficient available text to train quality language models. In an attempt to bridge this gap, we introduce the Norwegian Colossal Corpus (NCC), which comprises 49GB of cle…

NARC – Norwegian Anaphora Resolution Corpus

2022-10-01 · COLING (CRAC) 2022 10 · Petter Mæhlum, Dag Haug, Tollef Jørgensen, Andre Kåsen 외

We present the Norwegian Anaphora Resolution Corpus (NARC), the first publicly available corpus annotated with anaphoric relations between noun phrases for Norwegian. The paper describes the annotated data for 326 docume…

Relation

NorNE: Annotating Named Entities for Norwegian

2019-11-27 · LREC 2020 5 · Fredrik Jørgensen, Tobias Aasmoe, Anne-Stine Ruud Husevåg, Lilja Øvrelid 외

This paper presents NorNE, a manually annotated corpus of named entities which extends the annotation of the existing Norwegian Dependency Treebank. Comprising both of the official standards of written Norwegian (Bokm{\a…