paper-with-me

홈 › Papers

Storywrangler: A massive exploratorium for sociolinguistic, cultural, socioeconomic, and political timelines using Twitter

2020-07-25 · Thayer Alshaabi, Jane L. Adams, Michael V. Arnold, Joshua R. Minot, David R. Dewhurst, Andrew J. Reagan, Christopher M. Danforth, Peter Sheridan Dodds

In real-time, social media data strongly imprints world events, popular culture, and day-to-day conversations by millions of ordinary people at a scale that is scarcely conventionalized and recorded. Vitally, and absent from many standard corpora such as books and news archives, sharing and commenting mechanisms are native to social media platforms, enabling us to quantify social amplification (i.e., popularity) of trending storylines and contemporary cultural phenomena. Here, we describe Storywrangler, a natural language processing instrument designed to carry out an ongoing, day-scale curation of over 100 billion tweets containing roughly 1 trillion 1-grams from 2008 to 2021. For each day, we break tweets into unigrams, bigrams, and trigrams spanning over 100 languages. We track n-gram usage frequencies, and generate Zipf distributions, for words, hashtags, handles, numerals, symbols, and emojis. We make the data set available through an interactive time series viewer, and as downloadable time series and daily distributions. Although Storywrangler leverages Twitter data, our method of extracting and tracking dynamic changes of n-grams can be extended to any similar social media platform. We showcase a few examples of the many possible avenues of study we aim to enable including how social amplification can be visualized through 'contagiograms'. We also present some example case studies that bridge n-gram time series with disparate data sources to explore sociotechnical dynamics of famous individuals, box office success, and social unrest.

📄 PDF Abstract BibTeX arXiv:2007.12988

Code (6)

https://gitlab.com/compstorylab/contagiograms 공식 구현
https://gitlab.com/compstorylab/storywrangler 공식 구현
compstorylab/contagiograms
compstorylab/covid19ngrams
compstorylab/storywrangling
https://gitlab.com/compstorylab/covid19ngrams

Tasks

Cultural Vocal Bursts Intensity PredictionTime SeriesTime Series Analysis

Similar Papers 제목 키워드 기반

Socioeconomic Dependencies of Linguistic Patterns in Twitter: A Multivariate Analysis

2018-04-03 · Jacob Levy Abitbol, Márton Karsai, Jean-Philippe Magué, Jean-Pierre Chevrot 외

Our usage of language is not solely reliant on cognition but is arguably determined by myriad external factors leading to a global variability of linguistic patterns. This issue, which lies at the core of sociolinguistic…

Understanding Cultural Conflicts using Metaphors and Sociolinguistic Measures of Influence

2015-06-01 · WS 2015 6 · Samira Shaikh, Tomek Strzalkowski, Sarah Taylor, John Lien 외

Benchmarking Sociolinguistic Diversity in Swahili NLP: A Taxonomy-Guided Approach

2025-08-06 · Kezia Oketch, John P. Lalor, Ahmed Abbasi arxiv

We introduce the first taxonomy-guided evaluation of Swahili NLP, addressing gaps in sociolinguistic diversity. Drawing on health-related psychometric tasks, we collect a dataset of 2,170 free-text responses from Kenyan …

No Filter: Cultural and Socioeconomic Diversity in Contrastive Vision-Language Models

2024-05-22 · Angéline Pouget, Lucas Beyer, Emanuele Bugliarello, Xiao Wang 외

We study cultural and socioeconomic diversity in contrastive vision-language models (VLMs). Using a broad range of benchmark datasets and evaluation metrics, we bring to attention several important findings. First, the c…

Diversitygeo-localization

Crossing Borders Without Crossing Boundaries: How Sociolinguistic Awareness Can Optimize User Engagement with Localized Spanish AI Models Across Hispanophone Countries

2025-05-15 · Martin Capdevila, Esteban Villa Turek, Ellen Karina Chumbe Fernandez, Luis Felipe Polo Galvez 외

Large language models are, by definition, based on language. In an effort to underscore the critical need for regional localized models, this paper examines primary differences between variants of written Spanish across …