Creating a contemporary corpus of similes in Serbian by using natural language processing
Simile is a figure of speech that compares two things through the use of connection words, but where comparison is not intended to be taken literally. They are often used in everyday communication, but they are also a part of linguistic cultural heritage. In this paper we present a methodology for semi-automated collection of similes from the World Wide Web using text mining and machine learning techniques. We expanded an existing corpus by collecting 442 similes from the internet and adding them to the existing corpus collected by Vuk Stefanovic Karadzic that contained 333 similes. We, also, introduce crowdsourcing to the collection of figures of speech, which helped us to build corpus containing 787 unique similes.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
As Cool as a Cucumber: Towards a Corpus of Contemporary Similes in Serbian
Similes are natural language expressions used to compare unlikely things, where the comparison is not taken literally. They are often used in everyday communication and are an important part of cultural heritage. Having …
Analysis of Similes in Serbian Literary Texts (1860-1920) using computational methods
Similes are rhetorical figures which play an important role in literary texts. This paper presents a finite-state methodology developed for the description of adjectival similes, which enables their retrieval and annotat…
RetrievalSpecificityA Language-independent Model for Introducing a New Semantic Relation Between Adjectives and Nouns in a WordNet
The aim of this paper is to show a language-independent process of creating a new semantic relation between adjectives and nouns in wordnets. The existence of such a relation is expected to improve the detection of figur…
AttributeRelationSentiment AnalysisMachine Learning and Deep Neural Network-Based Lemmatization and Morphosyntactic Tagging for Serbian
The training of new tagger models for Serbian is primarily motivated by the enhancement of the existing tagset with the grammatical category of a gender. The harmonization of resources that were manually annotated within…
BIG-bench Machine LearningLemmatizationPOSPOS TaggingA Survey of Resources and Methods for Natural Language Processing of Serbian Language
The Serbian language is a Slavic language spoken by over 12 million speakers and well understood by over 15 million people. In the area of natural language processing, it can be considered a low-resourced language. Also,…
named-entity-recognitionNamed Entity Recognition