paper-with-me

Papers

Identifying Sentiments in Algerian Code-switched User-generated Comments

2020-05-01 · LREC 2020 5 · Wafia Adouane, Samia Touileb, Jean-Philippe Bernardy

We present in this paper our work on Algerian language, an under-resourced North African colloquial Arabic variety, for which we built a comparably large corpus of more than 36,000 code-switched user-generated comments annotated for sentiments. We opted for this data domain because Algerian is a colloquial language with no existing freely available corpora. Moreover, we compiled sentiment lexicons of positive and negative unigrams and bigrams reflecting the code-switches present in the language. We compare the performance of four models on the task of identifying sentiments, and the results indicate that a CNN model trained end-to-end fits better our unedited code-switched and unbalanced data across the predefined sentiment classes. Additionally, injecting the lexicons as background knowledge to the model boosts its performance on the minority class with a gain of 10.54 points on the F-score. The results of our experiments can be used as a baseline for future research for Algerian sentiment analysis.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Sentiment Analysis

Similar Papers 제목 키워드 기반

When is Multi-task Learning Beneficial for Low-Resource Noisy Code-switched User-generated Algerian Texts?

2020-05-01 · LREC 2020 5 · Wafia Adouane, Jean-Philippe Bernardy

We investigate when is it beneficial to simultaneously learn representations for several tasks, in low-resource settings. For this, we work with noisy user-generated texts in Algerian, a low-resource non-standardised Ara…

Data AugmentationMulti-Task Learningnamed-entity-recognitionNamed Entity Recognition+1

Towards Phone Number Recognition For Code Switched Algerian Dialect

2021-11-01 · ICNLSP 2021 11 · Khaled Lounnas, Mourad Abbas, Mohamed Lichouri

Normalising Non-standardised Orthography in Algerian Code-switched User-generated Data

2019-11-01 · WS 2019 11 · Wafia Adouane, Jean-Philippe Bernardy, Simon Dobnik

We work with Algerian, an under-resourced non-standardised Arabic variety, for which we compile a new parallel corpus consisting of user-generated textual data matched with normalised and corrected human annotations foll…

DecoderSemantic Textual SimilaritySpelling Correction

The interplay between language similarity and script on a novel multi-layer Algerian dialect corpus

2021-05-16 · Findings (ACL) 2021 8 · Samia Touileb, Jeremy Barnes

Recent years have seen a rise in interest for cross-lingual transfer between languages with similar typology, and between languages of various scripts. However, the interplay between language similarity and difference in…

Cross-Lingual TransferPart-Of-Speech TaggingSentiment Analysis

An End-to-End Hybrid Framework for Rumour Detection in Low-Resources Algerian Dialect

2026-06-11 · Dihia Lanasri, Fatima Benbarek arxiv

The rapid growth of social media has intensified the spread of rumours. This issue is more challenging in the Algerian context due to the informal and code-switched nature of dialectal content, the scarcity of annotated …

Rumour Detection