paper-with-me

홈 › Papers

Learning the Topic, Not the Language: How LLMs Classify Online Immigration Discourse Across Languages

2025-08-08 · Andrea Nasuto, Stefano Maria Iacus, Francisco Rowe, Devika Jain arxiv

Large language models (LLMs) offer new opportunities for scalable analysis of online discourse. Yet their use in multilingual social science research remains constrained by model size, cost and linguistic bias. We develop a lightweight, open-source LLM framework using fine-tuned LLaMA 3.2-3B models to classify immigration-related tweets across 13 languages. Unlike prior work relying on BERT style models or translation pipelines, we combine topic classification with stance detection and demonstrate that LLMs fine-tuned in just one or two languages can generalize topic understanding to unseen languages. Capturing ideological nuance, however, benefits from multilingual fine-tuning. Our approach corrects pretraining biases with minimal data from under-represented languages and avoids reliance on proprietary systems. With 26-168x faster inference and over 1000x cost savings compared to commercial LLMs, our method supports real-time analysis of billions of tweets. This scale-first framework enables inclusive, reproducible research on public attitudes across linguistic and cultural contexts.

📄 PDF Abstract BibTeX arXiv:2508.06435

Code (0)

등록된 구현이 없습니다.

Tasks

Stance Detection

Similar Papers 제목 키워드 기반

Integrating Large Language Models and Knowledge Graphs to Capture Political Viewpoints in News Media

2025-12-16 · Massimiliano Fadda, Enrico Motta, Francesco Osborne, Diego Reforgiato Recupero 외 arxiv

News sources play a central role in democratic societies by shaping political and social discourse through specific topics, viewpoints and voices. Understanding these dynamics is essential for assessing whether the media…

Knowledge Graphs

CIVICS: Building a Dataset for Examining Culturally-Informed Values in Large Language Models

2024-05-22 · Giada Pistilli, Alina Leidinger, Yacine Jernite, Atoosa Kasirzadeh 외

This paper introduces the "CIVICS: Culturally-Informed & Values-Inclusive Corpus for Societal impacts" dataset, designed to evaluate the social and cultural variation of Large Language Models (LLMs) across multiple langu…

Beyond Digital "Echo Chambers": The Role of Viewpoint Diversity in Political Discussion

2022-12-18 · Rishav Hada, Amir Ebrahimi Fard, Sarah Shugars, Federico Bianchi 외

Increasingly taking place in online spaces, modern political conversations are typically perceived to be unproductively affirming -- siloed in so called ``echo chambers'' of exclusively like-minded discussants. Yet, to d…

DiversityRecommendation Systems

Tracking Legislators’ Expressed Policy Agendas in Real Time

2022-01-18 · SocArXiv 2022 1 · Alexandra Siegel, David Laitin, Duncan Lawrence, Jeremy Weinstein 외

We develop a real-time scalable method to analyze strategic communication by political actors on salient policy issues through their tweets. Using word embeddings and supervised machine learning models, we classify legis…

Political Salient Issue Orientation DetectionText ClassificationWord Embeddings

Automated stance detection in complex topics and small languages: the challenging case of immigration in polarizing news media

2023-05-22 · Mark Mets, Andres Karjus, Indrek Ibrus, Maximilian Schich

Automated stance detection and related machine learning methods can provide useful insights for media monitoring and academic research. Many of these approaches require annotated training datasets, which limits their app…

Stance Detectiontext-classificationText Classification