paper-with-me

홈 › Papers

Detecting Basic Values in A Noisy Russian Social Media Text Data: A Multi-Stage Classification Framework

2026-03-19 · Maria Milkova, Maksim Rudnev arxiv

This study presents a multi-stage classification framework for detecting human values in noisy Russian language social media, validated on a random sample of 7.5 million public text posts. Drawing on Schwartz's theory of basic human values, we design a multi-stage pipeline that includes spam and nonpersonal content filtering, targeted selection of value relevant and politically relevant posts, LLM based annotation, and multi-label classification. Particular attention is given to verifying the quality of LLM annotations and model predictions against human experts. We treat human expert annotations not as ground truth but as an interpretative benchmark with its own uncertainty. To account for annotation subjectivity, we aggregate multiple LLM generated judgments into soft labels that reflect varying levels of agreement. These labels are then used to train transformer based models capable of predicting the probability of each of the ten basic values. The best performing model, XLM RoBERTa large, achieves an F1 macro of 0.83 and an F1 of 0.71 on held out test data. By treating value detection as a multi perspective interpretive task, where expert labels, GPT annotations, and model predictions represent coherent but not identical readings of the same texts, we show that the model generally aligns with human judgments but systematically overestimates the Openness to Change value domain. Empirically, the study reveals distinct patterns of value expression and their co-occurrence in Russian social networks, contributing to a broader research agenda on cultural variation, communicative framing, and value based interpretation in digital environments. All models are released publicly.

📄 PDF Abstract BibTeX arXiv:2603.18822

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-Label Classification

Similar Papers 제목 키워드 기반

Detecting value-expressive text posts in Russian social media

2023-12-14 · Maria Milkova, Maksim Rudnev, Lidia Okolskaya

Basic values are concepts or beliefs which pertain to desirable end-states and transcend specific situations. Studying personal values in social media can illuminate how and why societal values evolve especially when the…

Active LearningSpam detection

The methodology of Constructing the Large-Scale Dataset for Detecting Presuicidal and Anti-Suicidal Signals in Social Media Texts in Russian

2026-08-01 · Igor Buyanov, Darya Yaskova, Danil Serenko, Danil Shkereda 외 arxiv

The suicide is a terrifying act of a person who is misled by his own mental state. This problem arises across many countries. Sadly, Russia also has quite high number of persons who committed suicide. Luckily, a subset o…

BERT Implementation for Detecting Adverse Drug Effects Mentions in Russian

2020-12-01 · SMM4H (COLING) 2020 12 · Andrey Gusev, Anna Kuznetsova, Anna Polyanskaya, Egor Yatsishin

This paper describes a system developed for the Social Media Mining for Health 2020 shared task. Our team participated in the second subtask for Russian language creating a system to detect adverse drug reaction presence…

regression

Hidden Persuasion: Detecting Manipulative Narratives on Social Media During the 2022 Russian Invasion of Ukraine

2025-05-29 · Kateryna Akhynko, Oleksandr Kosovan, Mykola Trokhymovych

This paper presents one of the top-performing solutions to the UNLP 2025 Shared Task on Detecting Manipulation in Social Media. The task focuses on detecting and classifying rhetorical and stylistic manipulation techniqu…

Binary ClassificationClassificationLanguage ModelingLanguage Modelling

Detecting Troll Tweets in a Bilingual Corpus

2020-05-01 · LREC 2020 5 · Lin Miao, Mark Last, Marina Litvak

During the past several years, a large amount of troll accounts has emerged with efforts to manipulate public opinion on social network sites. They are often involved in spreading misinformation, fake news, and propagand…

Authorship VerificationFeature EngineeringGeneral ClassificationMisinformation+1