paper-with-me

홈 › Papers

RedDust: a Large Reusable Dataset of Reddit User Traits

2020-05-01 · LREC 2020 5 · Anna Tigunova, Paramita Mirza, Andrew Yates, Gerhard Weikum

Social media is a rich source of assertions about personal traits, such as {`}I am a doctor{''} or {`}my hobby is playing tennis{''}. Precisely identifying explicit assertions is difficult, though, because of the users{'} highly varied vocabulary and language expressions. Identifying personal traits from implicit assertions like I{'}ve been at work treating patients all day is even more challenging. This paper presents RedDust, a large-scale annotated resource for user profiling for over 300k Reddit users across five attributes: profession, hobby, family status, age,and gender. We construct RedDust using a diverse set of high-precision patterns and demonstrate its use as a resource for developing learning models to deal with implicit assertions. RedDust consists of users{'} personal traits, which are (attribute, value) pairs, along with users{'} post ids, which may be used to retrieve the posts from a publicly available crawl or from the Reddit API. We discuss the construction of the resource and show interesting statistics and insights into the data. We also compare different classifiers, which can be learned from RedDust. To the best of our knowledge, RedDust is the first annotated language resource about Reddit users at large scale. We envision further use cases of RedDust for providing background knowledge about user traits, to enhance personalized search and recommendation as well as conversational agents.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Attribute

Methods 이 논문이 사용한 방법론

AM 설명 없음

Similar Papers 제목 키워드 기반

Detection of Mental Health from Reddit via Deep Contextualized Representations

2020-11-01 · EMNLP (Louhi) 2020 11 · Zhengping Jiang, Sarah Ita Levitan, Jonathan Zomick, Julia Hirschberg

We address the problem of automatic detection of psychiatric disorders from the linguistic content of social media posts. We build a large scale dataset of Reddit posts from users with eight disorders and a control user …

ClassificationDiagnostic

TWeddit : A Dataset of Triggering Stories Predominantly Shared by Women on Reddit

2026-01-16 · Shirlene Rose Bandela, Sanjeev Parthasarathy, Vaibhav Garg arxiv

Warning: This paper may contain examples and topics that may be disturbing to some readers, especially survivors of miscarriage and sexual violence. People affected by abortion, miscarriage, or sexual violence often shar…

REALEDIT: Reddit Edits As a Large-scale Empirical Dataset for Image Transformations

2025-02-05 · CVPR 2025 1 · Peter Sushko, Ayana Bharadwaj, Zhi Yang Lim, Vasily Ilin 외

Existing image editing models struggle to meet real-world demands. Despite excelling in academic benchmarks, they have yet to be widely adopted for real user needs. Datasets that power these models use artificial edits, …

DeepFake DetectionFace Swapping

The Engage Corpus: A Social Media Dataset for Text-Based Recommender Systems

2022-06-01 · LREC 2022 6 · Daniel Cheng, Kyle Yan, Phillip Keung, Noah A. Smith

Social media platforms play an increasingly important role as forums for public discourse. Many platforms use recommendation algorithms that funnel users to online groups with the goal of maximizing user engagement, whic…

MisinformationRecommendation Systems

Comprehensive dataset of user-submitted articles with ideological and extreme bias from Reddit

2024-08-12 · Data in Brief 2024 8 · Kamalakkannan Ravi, Adan Ernesto Vela

Our study aims to collect data to understand ideological and extreme bias in text articles shared across various online communities, particularly focusing on the language used in subreddits associated with extremism and …

ArticlesHoldout SetNews ClassificationNews Recommendation+3