paper-with-me

Papers

The Arabic Parallel Gender Corpus 2.0: Extensions and Analyses

2021-10-18 · LREC 2022 6 · Bashar Alhafni, Nizar Habash, Houda Bouamor

Gender bias in natural language processing (NLP) applications, particularly machine translation, has been receiving increasing attention. Much of the research on this issue has focused on mitigating gender bias in English NLP models and systems. Addressing the problem in poorly resourced, and/or morphologically rich languages has lagged behind, largely due to the lack of datasets and resources. In this paper, we introduce a new corpus for gender identification and rewriting in contexts involving one or two target users (I and/or You) -- first and second grammatical persons with independent grammatical gender preferences. We focus on Arabic, a gender-marking morphologically rich language. The corpus has multiple parallel components: four combinations of 1st and 2nd person in feminine and masculine grammatical genders, as well as English, and English to Arabic machine translation output. This corpus expands on Habash et al. (2019)'s Arabic Parallel Gender Corpus (APGC v1.0) by adding second person targets as well as increasing the total number of sentences over 6.5 times, reaching over 590K words. Our new dataset will aid the research and development of gender identification, controlled text generation, and post-editing rewrite systems that could be used to personalize NLP applications and provide users with the correct outputs based on their grammatical gender preferences. We make the Arabic Parallel Gender Corpus (APGC v2.0) publicly available.

📄 PDF Abstract BibTeX arXiv:2110.09216

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationText GenerationTranslation

Similar Papers 제목 키워드 기반

ArabJobs: A Multinational Corpus of Arabic Job Ads

2025-09-26 · Mo El-Haj arxiv

ArabJobs is a publicly available corpus of Arabic job advertisements collected from Egypt, Jordan, Saudi Arabia, and the United Arab Emirates. Comprising over 8,500 postings and more than 550,000 words, the dataset captu…

Bias Detection

Gender Bias in Natural Language Processing Across Human Languages

2021-06-01 · NAACL (TrustNLP) 2021 6 · Abigail Matthews, Isabella Grasso, Christopher Mahoney, Yan Chen 외

Natural Language Processing (NLP) systems are at the heart of many critical automated decision-making systems making crucial recommendations about our future world. Gender bias in NLP has been well studied in English, bu…

Decision Making

Arap-Tweet: A Large Multi-Dialect Twitter Corpus for Gender, Age and Language Variety Identification

2018-08-23 · LREC 2018 5 · Wajdi Zaghouani, Anis Charfi

In this paper, we present Arap-Tweet, which is a large-scale and multi-dialectal corpus of Tweets from 11 regions and 16 countries in the Arab world representing the major Arabic dialectal varieties. To build this corpus…

Author Profiling

The SAMER Arabic Text Simplification Corpus

2024-04-29 · Bashar Alhafni, Reem Hazim, Juan Piñeros Liberato, Muhamed Al Khalil 외

We present the SAMER Corpus, the first manually annotated Arabic parallel corpus for text simplification targeting school-aged learners. Our corpus comprises texts of 159K words selected from 15 publicly available Arabic…

Text Simplification

A Multidialectal Parallel Corpus of Arabic

2014-05-01 · LREC 2014 5 · Houda Bouamor, Nizar Habash, Kemal Oflazer

The daily spoken variety of Arabic is often termed the colloquial or dialect form of Arabic. There are many Arabic dialects across the Arab World and within other Arabic speaking communities. These dialects vary widely f…

Dialect IdentificationMachine TranslationTranslation