paper-with-me

Papers

A Portuguese-Spanish Corpus Annotated for Subject Realization and Referentiality

2012-05-01 · LREC 2012 5 · Luz Rello, Iria Gayo

This paper presents a comparable corpus of Portuguese and Spanish consisting of legal and health texts. We describe the annotation of zero subject, impersonal constructions and explicit subjects in the corpus. We annotated 12,492 examples using a scheme that distinguishes between different linguistic levels (phonology, syntax, semantics, etc.) and present a taxonomy of instances on which annotators disagree. The high level of inter-annotator agreement (83{\%}-95{\%}) and the performance of learning algorithms trained on the corpus show that our corpus is a reliable and useful resource.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Coreference Resolution

Similar Papers 제목 키워드 기반

A New Annotated Portuguese/Spanish Corpus for the Multi-Sentence Compression Task

2018-05-01 · LREC 2018 5 · Elvys Linhares Pontes, Juan-Manuel Torres-Moreno, St{\'e}phane Huet, Andr{\'e}a Carneiro Linhares
Abstractive Text SummarizationQuestion AnsweringSentenceSentence Compression+1

TimeBankPT: A TimeML Annotated Corpus of Portuguese

2012-05-01 · LREC 2012 5 · Francisco Costa, Ant{\'o}nio Branco

In this paper, we introduce TimeBankPT, a TimeML annotated corpus of Portuguese. It has been produced by adapting an existing resource for English, namely the data used in the first TempEval challenge. TimeBankPT is the …

Machine TranslationTemporal Information Extractionvalid

Pay Attention when you Pay the Bills. A Multilingual Corpus with Dependency-based and Semantic Annotation of Collocations.

2019-07-01 · ACL 2019 7 · Marcos Garcia, Marcos Garc{\'\i}a Salido, Susana Sotelo, Estela Mosqueira 외

This paper presents a new multilingual corpus with semantic annotation of collocations in English, Portuguese, and Spanish. The whole resource contains 155k tokens and 1,526 collocations labeled in context. The annotated…

Natural Language UnderstandingText Generation

QTLeap WSD/NED Corpora: Semantic Annotation of Parallel Corpora in Six Languages

2016-05-01 · LREC 2016 5 · Arantxa Otegi, Nora Aranberri, Antonio Branco, Jan Haji{\v{c}} 외

This work presents parallel corpora automatically annotated with several NLP tools, including lemma and part-of-speech tagging, named-entity recognition and classification, named-entity disambiguation, word-sense disambi…

Cross-Lingual TransferEntity DisambiguationGeneral ClassificationLEMMA+7

Generating a Lexicon of Errors in Portuguese to Support an Error Identification System for Spanish Native Learners

2014-05-01 · LREC 2014 5 · Lianet Sep{\'u}lveda Torres, Magali Sanches Duran, S Alu{\'\i}sio, ra

Portuguese is a less resourced language in what concerns foreign language learning. Aiming to inform a module of a system designed to support scientific written production of Spanish native speakers learning Portuguese, …

Translation