paper-with-me

Papers

Introducing A large Tunisian Arabizi Dialectal Dataset for Sentiment Analysis

2021-04-01 · EACL (WANLP) 2021 4 · Chayma Fourati, Hatem Haddad, Abir Messaoudi, Moez BenHajhmida, Aymen Ben Elhaj Mabrouk, Malek Naski

On various Social Media platforms, people, tend to use the informal way to communicate, or write posts and comments: their local dialects. In Africa, more than 1500 dialects and languages exist. Particularly, Tunisians talk and write informally using Latin letters and numbers rather than Arabic ones. In this paper, we introduce a large common-crawl-based Tunisian Arabizi dialectal dataset dedicated for Sentiment Analysis. The dataset consists of a total of 100k comments (about movies, politic, sport, etc.) annotated manually by Tunisian native speakers as Positive, negative and Neutral. We evaluate our dataset on sentiment analysis task using the Bidirectional Encoder Representations from Transformers (BERT) as a contextual language model in its multilingual version (mBERT) as an embedding technique then combining mBERT with Convolutional Neural Network (CNN) as classifier. The dataset is publicly available.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingSentiment Analysis

Methods 이 논문이 사용한 방법론

mBERT mBERT

Similar Papers 제목 키워드 기반

TUNIZI: a Tunisian Arabizi sentiment analysis Dataset

2020-04-29 · Chayma Fourati, Abir Messaoudi, Hatem Haddad

On social media, Arabic people tend to express themselves in their own local dialects. More particularly, Tunisians use the informal way called "Tunisian Arabizi". Analytical studies seek to explore and recognize online …

MarketingSentiment Analysis

Multi-Task Sequence Prediction For Tunisian Arabizi Multi-Level Annotation

2020-11-10 · COLING (WANLP) 2020 12 · Elisa Gugliotta, Marco Dinarelli, Olivier Kraif

In this paper we propose a multi-task sequence prediction system, based on recurrent neural networks and used to annotate on multiple levels an Arabizi Tunisian corpus. The annotation performed are text classification, t…

POSPOS Taggingtext-classificationText Classification

TArC: Tunisian Arabish Corpus First complete release

2022-07-11 · Elisa Gugliotta, Marco Dinarelli

In this paper we present the final result of a project on Tunisian Arabic encoded in Arabizi, the Latin-based writing system for digital conversations. The project led to the creation of two integrated and independent re…

LemmatizationPOSPOS TaggingTransliteration

A Conventional Orthography for Tunisian Arabic

2014-05-01 · LREC 2014 5 · In{\`e}s Zribi, Rahma Boujelbane, Abir Masmoudi, Mariem ellouze 외

Tunisian Arabic is a dialect of the Arabic language spoken in Tunisia. Tunisian Arabic is an under-resourced language. It has neither a standard orthography nor large collections of written text and dictionaries. Actuall…

Language ModellingMachine TranslationSpeech RecognitionSpeech Synthesis+1

SenZi: A Sentiment Analysis Lexicon for the Latinised Arabic (Arabizi)

2019-09-01 · RANLP 2019 9 · Taha Tobaili, Fern, Miriam ez, Harith Alani 외

Arabizi is an informal written form of dialectal Arabic transcribed in Latin alphanumeric characters. It has a proven popularity on chat platforms and social media, yet it suffers from a severe lack of natural language p…

Sentiment AnalysisWord Embeddings