paper-with-me

홈 › Papers

More Romanian word embeddings from the RETEROM project

2021-11-21 · Vasile Păiş, Dan Tufiş

Automatically learned vector representations of words, also known as "word embeddings", are becoming a basic building block for more and more natural language processing algorithms. There are different ways and tools for constructing word embeddings. Most of the approaches rely on raw texts, the construction items being the word occurrences and/or letter n-grams. More elaborated research is using additional linguistic features extracted after text preprocessing. Morphology is clearly served by vector representations constructed from raw texts and letter n-grams. Syntax and semantics studies may profit more from the vector representations constructed with additional features such as lemma, part-of-speech, syntactic or semantic dependants associated with each word. One of the key objectives of the ReTeRom project is the development of advanced technologies for Romanian natural language processing, including morphological, syntactic and semantic analysis of text. As such, we plan to develop an open-access large library of ready-to-use word embeddings sets, each set being characterized by different parameters: used features (wordforms, letter n-grams, lemmas, POSes etc.), vector lengths, window/context size and frequency thresholds. To this end, the previously created sets of word embeddings (based on word occurrences) on the CoRoLa corpus (P\u{a}i\c{s} and Tufi\c{s}, 2018) are and will be further augmented with new representations learned from the same corpus by using specific features such as lemmas and parts of speech. Furthermore, in order to better understand and explore the vectors, graphical representations will be available by customized interfaces.

📄 PDF Abstract BibTeX arXiv:2111.10750

Code (0)

등록된 구현이 없습니다.

Tasks

LEMMAWord Embeddings

Similar Papers 제목 키워드 기반

CoBiLiRo: A Research Platform for Bimodal Corpora

2020-05-01 · LREC 2020 5 · Dan Cristea, Ionu{\textcommabelow{t}} Pistol, {\textcommabelow{S}}erban Boghiu, Anca-Diana Bibiri 외

This paper describes the on-going work carried out within the CoBiLiRo (Bimodal Corpus for Romanian Language) research project, part of ReTeRom (Resources and Technologies for Developing Human-Machine Interfaces in Roman…

speech-recognitionSpeech Recognition

Clustering Word Embeddings with Self-Organizing Maps. Application on LaRoSeDa - A Large Romanian Sentiment Data Set

2021-04-01 · EACL 2021 2 · Anca Tache, Gaman Mihaela, Radu Tudor Ionescu

Romanian is one of the understudied languages in computational linguistics, with few resources available for the development of natural language processing tools. In this paper, we introduce LaRoSeDa, a Large Romanian Se…

ClusteringSentiment AnalysisSentiment ClassificationText Categorization+1

Clustering Word Embeddings with Self-Organizing Maps. Application on LaRoSeDa -- A Large Romanian Sentiment Data Set

2021-01-11 · Anca Maria Tache, Mihaela Gaman, Radu Tudor Ionescu

Romanian is one of the understudied languages in computational linguistics, with few resources available for the development of natural language processing tools. In this paper, we introduce LaRoSeDa, a Large Romanian Se…

ClusteringSentiment AnalysisSentiment ClassificationText Categorization+1

Named Entity Recognition in the Romanian Legal Domain

2021-11-01 · EMNLP (NLLP) 2021 11 · Vasile Pais, Maria Mitrofan, Carol Luca Gasan, Vlad Coneschi 외

Recognition of named entities present in text is an important step towards information extraction and natural language understanding. This work presents a named entity recognition system for the Romanian legal domain. Th…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Natural Language Understanding+1

Use Case: Romanian Language Resources in the LOD Paradigm

2022-06-01 · LDL (ACL) 2022 6 · Verginica Barbu Mititelu, Elena Irimia, Vasile Pais, Andrei-Marius Avram 외

In this paper, we report on (i) the conversion of Romanian language resources to the Linked Open Data specifications and requirements, on (ii) their publication and (iii) interlinking with other language resources (for R…

Word Embeddings