paper-with-me

홈 › Papers

BERTuit: Understanding Spanish language in Twitter through a native transformer

2022-04-07 · Javier Huertas-Tato, Alejandro Martin, David Camacho

The appearance of complex attention-based language models such as BERT, Roberta or GPT-3 has allowed to address highly complex tasks in a plethora of scenarios. However, when applied to specific domains, these models encounter considerable difficulties. This is the case of Social Networks such as Twitter, an ever-changing stream of information written with informal and complex language, where each message requires careful evaluation to be understood even by humans given the important role that context plays. Addressing tasks in this domain through Natural Language Processing involves severe challenges. When powerful state-of-the-art multilingual language models are applied to this scenario, language specific nuances use to get lost in translation. To face these challenges we present \textbf{BERTuit}, the larger transformer proposed so far for Spanish language, pre-trained on a massive dataset of 230M Spanish tweets using RoBERTa optimization. Our motivation is to provide a powerful resource to better understand Spanish Twitter and to be used on applications focused on this social network, with special emphasis on solutions devoted to tackle the spreading of misinformation in this platform. BERTuit is evaluated on several tasks and compared against M-BERT, XLM-RoBERTa and XLM-T, very competitive multilingual transformers. The utility of our approach is shown with applications, in this case: a zero-shot methodology to visualize groups of hoaxes and profiling authors spreading disinformation. Misinformation spreads wildly on platforms such as Twitter in languages other than English, meaning performance of transformers may suffer when transferred outside English speaking communities.

📄 PDF Abstract BibTeX arXiv:2204.03465

Code (0)

등록된 구현이 없습니다.

Tasks

Misinformation

Methods 이 논문이 사용한 방법론

15 Ways to Contact How can i speak to someone at Delta Airlines 설명 없음
Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Cosine Annealing Cosine Annealing is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before…
Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Adam 설명 없음
Residual Connection 설명 없음

Similar Papers 제목 키워드 기반

Sentiment Analysis of Spanish Political Party Tweets Using Pre-trained Language Models

2024-11-07 · Chuqiao Song, Shunzhang Chen, Xinyi Cai, Hao Chen

Title: Sentiment Analysis of Spanish Political Party Communications on Twitter Using Pre-trained Language Models Authors: Chuqiao Song, Shunzhang Chen, Xinyi Cai, Hao Chen Comments: 21 pages, 6 figures Abstract: This stu…

Sentiment Analysis

RoBERTuito: a pre-trained language model for social media text in Spanish

2021-11-18 · LREC 2022 6 · Juan Manuel Pérez, Damián A. Furman, Laura Alonso Alemany, Franco Luque

Since BERT appeared, Transformer language models and transfer learning have become state-of-the-art for Natural Language Understanding tasks. Recently, some works geared towards pre-training specially-crafted models for …

Language ModelingLanguage ModellingNatural Language UnderstandingTransfer Learning

Regionalized models for Spanish language variations based on Twitter

2021-10-12 · Eric S. Tellez, Daniela Moctezuma, Sabino Miranda, Mario Graff 외

Spanish is one of the most spoken languages in the globe, but not necessarily Spanish is written and spoken in the same way in different countries. Understanding local language variations can help to improve model perfor…

Word Embeddings

Information Privacy Opinions on Twitter: A Cross-Language Study

2019-12-05 · Felipe González, Andrea Figueroa, Claudia López, Cecilia Aragón

The Cambridge Analytica scandal triggered a conversation on Twitter about data practices and their implications. Our research proposes to leverage this conversation to extend the understanding of how information privacy …

Crowdsourcing Dialect Characterization through Twitter

2014-07-26 · Bruno Gonçalves, David Sánchez

We perform a large-scale analysis of language diatopic variation using geotagged microblogging datasets. By collecting all Twitter messages written in Spanish over more than two years, we build a corpus from which a care…