RusLICA: A Russian-Language Platform for Automated Linguistic Inquiry and Category Analysis
Defining psycholinguistic characteristics in written texts is a task gaining increasing attention from researchers. One of the most widely used tools in the current field is Linguistic Inquiry and Word Count (LIWC) that originally was developed to analyze English texts and translated into multiple languages. Our approach offers the adaptation of LIWC methodology for the Russian language, considering its grammatical and cultural specificities. The suggested approach comprises 96 categories, integrating syntactic, morphological, lexical, general statistical features, and results of predictions obtained using pre-trained language models (LMs) for text analysis. Rather than applying direct translation to existing thesauri, we built the dictionary specifically for the Russian language based on the content from several lexicographic resources, semantic dictionaries and corpora. The paper describes the process of mapping lemmas to 42 psycholinguistic categories and the implementation of the analyzer as part of RusLICA web service.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Deception Detection for the Russian Language: Lexical and Syntactic Parameters
The field of automated deception detection in written texts is methodologically challenging. Different linguistic levels (lexics, syntax and semantics) are basically used for different types of English texts to reveal if…
Binary text classificationDeception DetectionInformation RetrievalPOS+2Automated WordNet Construction Using Word Embeddings
We present a fully unsupervised method for automated construction of WordNets based upon recent advances in distributional representations of sentences and word-senses combined with readily available machine translation …
Information RetrievalMachine TranslationTranslationWord Embeddings+1Detecting Spelling and Grammatical Anomalies in Russian Poetry Texts
The quality of natural language texts in fine-tuning datasets plays a critical role in the performance of generative models, particularly in computational creativity tasks such as poem or song lyric generation. Fluency d…
Anomaly DetectionGrammatical Error DetectionSentenceGenerating Conceptual Metaphors from Proposition Stores
Contemporary research on computational processing of linguistic metaphors is divided into two main branches: metaphor recognition and metaphor interpretation. We take a different line of research and present an automated…
Class-based LSTM Russian Language Model with Linguistic Information
In the paper, we present class-based LSTM Russian language models (LMs) with classes generated with the use of both word frequency and linguistic information data, obtained with the help of the {``}VisualSynan{''} softwa…
Language ModelingLanguage Modellingspeech-recognitionSpeech Recognition