paper-with-me

홈 › Papers

As Cool as a Cucumber: Towards a Corpus of Contemporary Similes in Serbian

2016-05-20 · Nikola Milosevic, Goran Nenadic

Similes are natural language expressions used to compare unlikely things, where the comparison is not taken literally. They are often used in everyday communication and are an important part of cultural heritage. Having an up-to-date corpus of similes is challenging, as they are constantly coined and/or adapted to the contemporary times. In this paper we present a methodology for semi-automated collection of similes from the world wide web using text mining techniques. We expanded an existing corpus of traditional similes (containing 333 similes) by collecting 446 additional expressions. We, also, explore how crowdsourcing can be used to extract and curate new similes.

📄 PDF Abstract BibTeX arXiv:1605.06319

Code (1)

nikolamilosevic86/SerbianComparisonExtractor 공식 구현

Similar Papers 제목 키워드 기반

Creating a contemporary corpus of similes in Serbian by using natural language processing

2018-11-22 · Nikola Milosevic, Goran Nenadic

Simile is a figure of speech that compares two things through the use of connection words, but where comparison is not intended to be taken literally. They are often used in everyday communication, but they are also a pa…

Analysis of Similes in Serbian Literary Texts (1860-1920) using computational methods

2020-09-01 · CLIB 2020 9 · Cvetana Krstev, Jelena Jaćimović, Duško Vitas

Similes are rhetorical figures which play an important role in literary texts. This paper presents a finite-state methodology developed for the description of adjectival similes, which enables their retrieval and annotat…

RetrievalSpecificity

Machine Learning and Deep Neural Network-Based Lemmatization and Morphosyntactic Tagging for Serbian

2020-05-01 · LREC 2020 5 · Ranka Stankovic, {\v{S}}, Branislava rih, Cvetana Krstev 외

The training of new tagger models for Serbian is primarily motivated by the enhancement of the existing tagset with the grammatical category of a gender. The harmonization of resources that were manually annotated within…

BIG-bench Machine LearningLemmatizationPOSPOS Tagging

A Language-independent Model for Introducing a New Semantic Relation Between Adjectives and Nouns in a WordNet

2016-01-01 · GWC 2016 1 · Miljana Mladenović, Jelena Mitrović, Cvetana Krstev

The aim of this paper is to show a language-independent process of creating a new semantic relation between adjectives and nouns in wordnets. The existence of such a relation is expected to improve the detection of figur…

AttributeRelationSentiment Analysis

TALC-sef A Manually-Revised POS-TAgged Literary Corpus in Serbian, English and French

2014-05-01 · LREC 2014 5 · Antonio Balvet, Dejan Stosic, Aleks Miletic, ra

In this paper, we present a parallel literary corpus for Serbian, English and French, the TALC-sef corpus. The corpus includes a manually-revised pos-tagged reference Serbian corpus of over 150,000 words. The initial obj…

POSPOS Tagging