paper-with-me

홈 › Papers

Sociolinguistic Corpus of WhatsApp Chats in Spanish among College Students

2018-07-01 · WS 2018 7 · Alej Dorantes, ro, Gerardo Sierra, Tlauhlia Yam{\'\i}n Donohue P{\'e}rez, Gemma Bel-Enguix, M{\'o}nica Jasso Rosales

This work presents the Sociolinguistic Corpus of WhatsApp Chats in Spanish among College Students, a corpus of raw data for general use. Its purpose is to offer data for the study of of language and interactions via Instant Messaging (IM) among bachelors. Our paper consists of an overview of both the corpus{'}s content and demographic metadata. Furthermore, it presents the current research being conducted with it {---}namely parenthetical expressions, orality traits, and code-switching. This work also includes a brief outline of similar corpora and recent studies in the field of IM.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Modeling Topics and Sociolinguistic Variation in Code-Switched Discourse: Insights from Spanish-English and Spanish-Guaraní

2025-12-03 · Nemika Tyagi, Nelvin Licona Guevara, Olga Kellert arxiv

This study presents an LLM-assisted annotation pipeline for the sociolinguistic and topical analysis of bilingual discourse in two typologically distinct contexts: Spanish-English and Spanish-Guaraní. Using large languag…

Creating a WhatsApp Dataset to Study Pre-teen Cyberbullying

2018-10-01 · WS 2018 10 · Rachele Sprugnoli, Stefano Menini, Sara Tonelli, Filippo Oncini 외

Although WhatsApp is used by teenagers as one major channel of cyberbullying, such interactions remain invisible due to the app privacy policies that do not allow ex-post data collection. Indeed, most of the information …

Crossing Borders Without Crossing Boundaries: How Sociolinguistic Awareness Can Optimize User Engagement with Localized Spanish AI Models Across Hispanophone Countries

2025-05-15 · Martin Capdevila, Esteban Villa Turek, Ellen Karina Chumbe Fernandez, Luis Felipe Polo Galvez 외

Large language models are, by definition, based on language. In an effort to underscore the critical need for regional localized models, this paper examines primary differences between variants of written Spanish across …

A Linguistic Annotation Framework to Study Interactions in Multilingual Healthcare Conversational Forums

2021-11-01 · EMNLP (LAW, DMR) 2021 11 · Ishani Mondal, Kalika Bali, Mohit Jain, Monojit Choudhury 외

In recent years, remote digital healthcare using online chats has gained momentum, especially in the Global South. Though prior work has studied interaction patterns in online (health) forums, such as TalkLife, Reddit an…

Disentangling Codemixing in Chats: The NUS ABC Codemixed Corpus

2025-05-31 · Svetlana Churina, Akshat Gupta, Insyirah Mujtahid, Kokil Jaidka

Code-mixing involves the seamless integration of linguistic elements from multiple languages within a single discourse, reflecting natural multilingual communication patterns. Despite its prominence in informal interacti…