paper-with-me

Papers

Challenges in Creating a Representative Corpus of Romanian Micro-Blogging Text

2022-06-01 · CMLC (LREC) 2022 6 · Vasile Pais, Maria Mitrofan, Verginica Barbu Mititelu, Elena Irimia, Roxana Micu, Carol Luca Gasan

Following the successful creation of a national representative corpus of contemporary Romanian language, we turned our attention to the social media text, as present in micro-blogging platforms. In this paper, we present the current activities as well as the challenges faced when trying to apply existing tools (for both annotation and indexing) to a Romanian language micro-blogging corpus. These challenges are encountered at all annotation levels, including tokenization, and at the indexing stage. We consider that existing tools for Romanian language processing must be adapted to recognize features such as emoticons, emojis, hashtags, unusual abbreviations, elongated words (commonly used for emphasis in micro-blogging), multiple words joined together (within oroutside hashtags), and code-mixed text.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Evaluating the Wordnet and CoRoLa-based Word Embedding Vectors for Romanian as Resources in the Task of Microworlds Lexicon Expansion

2019-07-01 · GWC 2019 7 · Elena Irimia, Maria Mitrofan, Verginica Mititelu

Within a larger frame of facilitating human-robot interaction, we present here the creation of a core vocabulary to be learned by a robot. It is extracted from two tokenised and lemmatized scenarios pertaining to two ima…

Resources in Underrepresented Languages: Building a Representative Romanian Corpus

2020-05-01 · LREC 2020 5 · Ludmila Midrigan - Ciochina, Victoria Boyd, Lucila Sanchez-Ortega, Diana Malancea{\_}Malac 외

The effort in the field of Linguistics to develop theories that aim to explain language-dependent effects on language processing is greatly facilitated by the availability of reliable resources representing different lan…

A Diachronic Corpus for Romanian (RoDia)

2017-09-01 · RANLP 2017 9 · Ludmila Malahov, C{\u{a}}t{\u{a}}lina M{\u{a}}r{\u{a}}nduc, Alex Colesnicov, ru

This paper describes a Romanian Dependency Treebank, built at the Al. I. Cuza University (UAIC), and a special OCR techniques used to build it. The corpus has rich morphological and syntactic annotation. There are few an…

Information RetrievalOptical Character Recognition (OCR)Question Answering

Romanian micro-blogging named entity recognition including health-related entities

2022-10-01 · SMM4H (COLING) 2022 10 · Vasile Pais, Verginica Barbu Mititelu, Elena Irimia, Maria Mitrofan 외

This paper introduces a manually annotated dataset for named entity recognition (NER) in micro-blogging text for Romanian language. It contains gold annotations for 9 entity classes and expressions: persons, locations, o…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+1

CoRoLa --- The Reference Corpus of Contemporary Romanian Language

2014-05-01 · LREC 2014 5 · Verginica Barbu Mititelu, Elena Irimia, Dan Tufi{\textcommabelow{s}}

We present the project of creating CoRoLa, a reference corpus of contemporary Romanian (from 1945 onwards). In the international context, the project finds its place among the initiatives of gathering huge collections of…

LemmatizationSentence