paper-with-me

Papers

Using Noisy Self-Reports to Predict Twitter User Demographics

2020-05-01 · NAACL (SocialNLP) 2021 6 · Zach Wood-Doughty, Paiheng Xu, Xiao Liu, Mark Dredze

Computational social science studies often contextualize content analysis within standard demographics. Since demographics are unavailable on many social media platforms (e.g. Twitter) numerous studies have inferred demographics automatically. Despite many studies presenting proof of concept inference of race and ethnicity, training of practical systems remains elusive since there are few annotated datasets. Existing datasets are small, inaccurate, or fail to cover the four most common racial and ethnic groups in the United States. We present a method to identify self-reports of race and ethnicity from Twitter profile descriptions. Despite errors inherent in automated supervision, we produce models with good performance when measured on gold standard self-report survey data. The result is a reproducible method for creating large-scale training resources for race and ethnicity.

📄 PDF Abstract BibTeX arXiv:2005.00635

Code (1)

https://bitbucket.org/mdredze/demographer 공식 구현 tf

Similar Papers 제목 키워드 기반

Using Noisy Self-Reports to Predict Twitter User Demographics

2019-11-18 · Anonymous

Computational social science studies often contextualize content analysis within standard demographics. Since demographic attributes are unavailable on many social media platforms, such as Twitter, numerous studies have…

Inferring Fine-grained Details on User Activities and Home Location from Social Media: Detecting Drinking-While-Tweeting Patterns in Communities

2016-03-10 · Nabil Hossain, Tianran Hu, Roghayeh Feizi, Ann Marie White 외

Nearly all previous work on geo-locating latent states and activities from social media confounds general discussions about activities, self-reports of users participating in those activities at times in the past or futu…

Evaluating the effectiveness of Phishing Reports on Twitter

2021-11-13 · Sayak Saha Roy, Unique Karanjit, Shirin Nilizadeh

Phishing attacks are an increasingly potent web-based threat, with nearly 1.5 million websites created on a monthly basis. In this work, we present the first study towards identifying such attacks through phishing report…

TwHIN-BERT: A Socially-Enriched Pre-trained Language Model for Multilingual Tweet Representations at Twitter

2022-09-15 · Xinyang Zhang, Yury Malkov, Omar Florez, Serim Park 외

Pre-trained language models (PLMs) are fundamental for natural language processing applications. Most existing PLMs are not tailored to the noisy user-generated text on social media, and the pre-training does not factor …

Language ModelingLanguage Modelling

Multi-View Active Learning for Short Text Classification in User-Generated Data

2021-12-05 · Payam Karisani, Negin Karisani, Li Xiong

Mining user-generated data often suffers from the lack of enough labeled data, short document lengths, and the informal user language. In this paper, we propose a novel active learning model to overcome these obstacles i…

Active LearningLanguage Modellingtext-classificationText Classification