paper-with-me

홈 › Papers

The Secret Lives of Names? Name Embeddings from Social Media

2019-05-12 · Junting Ye, Steven Skiena

Your name tells a lot about you: your gender, ethnicity and so on. It has been shown that name embeddings are more effective in representing names than traditional substring features. However, our previous name embedding model is trained on private email data and are not publicly accessible. In this paper, we explore learning name embeddings from public Twitter data. We argue that Twitter embeddings have two key advantages: \textit{(i)} they can and will be publicly released to support research community. \textit{(ii)} even with a smaller training corpus, Twitter embeddings achieve similar performances on multiple tasks comparing to email embeddings. As a test case to show the power of name embeddings, we investigate the modeling of lifespans. We find it interesting that adding name embeddings can further improve the performances of models using demographic features, which are traditionally used for lifespan modeling. Through residual analysis, we observe that fine-grained groups (potentially reflecting socioeconomic status) are the latent contributing factors encoded in name embeddings. These were previously hidden to demographic models, and may help to enhance the predictive power of a wide class of research studies.

📄 PDF Abstract BibTeX arXiv:1905.04799

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Regional Negative Bias in Word Embeddings Predicts Racial Animus--but only via Name Frequency

2022-01-20 · Austin Van Loon, Salvatore Giorgi, Robb Willer, Johannes Eichstaedt

The word embedding association test (WEAT) is an important method for measuring linguistic biases against social groups such as ethnic minorities in large text corpora. It does so by comparing the semantic relatedness of…

AttributeWord Embeddings

On the Influence of Gender and Race in Romantic Relationship Prediction from Large Language Models

2024-10-05 · Abhilasha Sancheti, Haozhe An, Rachel Rudinger

We study the presence of heteronormative biases and prejudice against interracial romantic relationships in large language models by performing controlled name-replacement experiments for the task of relationship predict…

What are the biases in my word embedding?

2018-12-20 · Nathaniel Swinger, Maria De-Arteaga, Neil Thomas Heffernan IV, Mark DM Leiserson 외

This paper presents an algorithm for enumerating biases in word embeddings. The algorithm exposes a large number of offensive associations related to sensitive features such as race and gender on publicly available embed…

Word Embeddings

SA2SL: From Aspect-Based Sentiment Analysis to Social Listening System for Business Intelligence

2021-05-31 · Luong Luc Phan, Phuc Huynh Pham, Kim Thi-Thanh Nguyen, Tham Thi Nguyen 외

In this paper, we present a process of building a social listening system based on aspect-based sentiment analysis in Vietnamese from creating a dataset to building a real application. Firstly, we create UIT-ViSFD, a Vie…

Aspect-Based Sentiment AnalysisAspect-Based Sentiment Analysis (ABSA)General ClassificationSentiment Analysis+4

ViCGCN: Graph Convolutional Network with Contextualized Language Models for Social Media Mining in Vietnamese

2023-09-06 · Chau-Thang Phan, Quoc-Nam Nguyen, Chi-Thanh Dang, Trong-Hop Do 외

Social media processing is a fundamental task in natural language processing with numerous applications. As Vietnamese social media and information science have grown rapidly, the necessity of information-based mining on…

Language Modellingtext-classificationText ClassificationVietnamese Social Media Text Processing