The Engage Corpus: A Social Media Dataset for Text-Based Recommender Systems
Social media platforms play an increasingly important role as forums for public discourse. Many platforms use recommendation algorithms that funnel users to online groups with the goal of maximizing user engagement, which many commentators have pointed to as a source of polarization and misinformation. Understanding the role of NLP in recommender systems is an interesting research area, given the role that social media has played in world events. However, there are few standardized resources which researchers can use to build models that predict engagement with online groups on social media; each research group constructs datasets from scratch without releasing their version for reuse. In this work, we present a dataset drawn from posts and comments on the online message board Reddit. We develop baseline models for recommending subreddits to users, given the user’s post and comment history. We also study the behavior of our recommender models on subreddits that were banned in June 2020 as part of Reddit’s efforts to stop the dissemination of hate speech.
Code (1)
Tasks
MisinformationRecommendation SystemsSimilar Papers 제목 키워드 기반
MultiMediate'23: Engagement Estimation and Bodily Behaviour Recognition in Social Interactions
Automatic analysis of human behaviour is a fundamental prerequisite for the creation of machines that can effectively interact with- and support humans in social interactions. In MultiMediate'23, we address two key human…
How-to Present News on Social Media: A Causal Analysis of Editing News Headlines for Boosting User Engagement
To reach a broader audience and optimize traffic toward news articles, media outlets commonly run social media accounts and share their content with a short text summary. Despite its importance of writing a compelling me…
ArticlesCausal InferencecounterfactualEnhancing Rumor Detection Methods with Propagation Structure Infused Language Model
Pretrained Language Models (PLMs) have excelled in various Natural Language Processing tasks, benefiting from large-scale pretraining and self-attention mechanism's ability to capture long-range dependencies. However, th…
DREAMS: A Social Exchange Theory-Informed Modeling of Misinformation Engagement on Social Media
Social media engagement prediction is a central challenge in computational social science, particularly for understanding how users interact with misinformation. Existing approaches often treat engagement as a homogeneou…
JobArabi: An Arabic Corpus and Analysis of Job Announcements from Social Media
This paper introduces JobArabi, a large-scale corpus of Arabic job announcements collected from social media between January 2024 and October 2025. The dataset contains 20,528 public posts from X and captures more than t…