paper-with-me

Papers

Learning about Spanish dialects through Twitter

2015-11-16 · Bruno Gonçalves, David Sánchez

This paper maps the large-scale variation of the Spanish language by employing a corpus based on geographically tagged Twitter messages. Lexical dialects are extracted from an analysis of variants of tens of concepts. The resulting maps show linguistic variation on an unprecedented scale across the globe. We discuss the properties of the main dialects within a machine learning approach and find that varieties spoken in urban areas have an international character in contrast to country areas where dialects show a more regional uniformity.

📄 PDF Abstract BibTeX arXiv:1511.04970

Code (0)

등록된 구현이 없습니다.

Tasks

BIG-bench Machine Learning

Similar Papers 제목 키워드 기반

Crowdsourcing Dialect Characterization through Twitter

2014-07-26 · Bruno Gonçalves, David Sánchez

We perform a large-scale analysis of language diatopic variation using geotagged microblogging datasets. By collecting all Twitter messages written in Spanish over more than two years, we build a corpus from which a care…

A Python Library for Exploratory Data Analysis on Twitter Data based on Tokens and Aggregated Origin-Destination Information

2020-09-03 · Mario Graff, Daniela Moctezuma, Sabino Miranda-Jiménez, Eric S. Tellez

Twitter is perhaps the social media more amenable for research. It requires only a few steps to obtain information, and there are plenty of libraries that can help in this regard. Nonetheless, knowing whether a particula…

Jojajovai: A Parallel Guarani-Spanish Corpus for MT Benchmarking

2022-06-01 · LREC 2022 6 · Luis Chiruzzo, Santiago Góngora, Aldo Alvarez, Gustavo Giménez-Lugo 외

This work presents a parallel corpus of Guarani-Spanish text aligned at sentence level. The corpus contains about 30,000 sentence pairs, and is structured as a collection of subsets from different sources, further split …

BenchmarkingSentenceTranslation

Dialect and Gender Bias in YouTube's Spanish Captioning System

2026-02-27 · Iris Dania Jimenez, Christoph Kern arxiv

Spanish is the official language of twenty-one countries and is spoken by over 441 million people. Naturally, there are many variations in how Spanish is spoken across these countries. Media platforms such as YouTube rel…

Speech Recognition

Information Privacy Opinions on Twitter: A Cross-Language Study

2019-12-05 · Felipe González, Andrea Figueroa, Claudia López, Cecilia Aragón

The Cambridge Analytica scandal triggered a conversation on Twitter about data practices and their implications. Our research proposes to leverage this conversation to extend the understanding of how information privacy …