paper-with-me

Papers

Can Crowdsourcing be used for Effective Annotation of Arabic?

2014-05-01 · LREC 2014 5 · Wajdi Zaghouani, Kais Dukes

Crowdsourcing has been used recently as an alternative to traditional costly annotation by many natural language processing groups. In this paper, we explore the use of Amazon Mechanical Turk (AMT) in order to assess the feasibility of using AMT workers (also known as Turkers) to perform linguistic annotation of Arabic. We used a gold standard data set taken from the Quran corpus project annotated with part-of-speech and morphological information. An Arabic language qualification test was used to filter out potential non-qualified participants. Two experiments were performed, a part-of-speech tagging task in where the annotators were asked to choose a correct word-category from a multiple choice list and case ending identification task. The results obtained so far showed that annotating Arabic grammatical case is harder than POS tagging, and crowdsourcing for Arabic linguistic annotation requiring expert annotators could be not as effective as other crowdsourcing experiments requiring less expertise and qualifications.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Entity ResolutionMultiple-choiceNatural Language InferencePart-Of-Speech TaggingPOSPOS TaggingWord Sense Disambiguation

Similar Papers 제목 키워드 기반

Quranic Audio Dataset: Crowdsourced and Labeled Recitation from Non-Arabic Speakers

2024-05-04 · Raghad Salameh, Mohamad Al Mdfaa, Nursultan Askarbekuly, Manuel Mazzara

This paper addresses the challenge of learning to recite the Quran for non-Arabic speakers. We explore the possibility of crowdsourcing a carefully annotated Quranic dataset, on top of which AI models can be built to sim…

Annotating Targets of Opinions in Arabic using Crowdsourcing

2015-07-01 · WS 2015 7 · Noura Farra, Kathy Mckeown, Nizar Habash
Fine-Grained Opinion AnalysisSubjectivity Analysis

Best Practices for Crowdsourcing Dialectal Arabic Speech Transcription

2015-07-01 · WS 2015 7 · Samantha Wray, Hamdy Mubarak, Ahmed Ali
Speech Recognition

A Crowdsourcing-based Approach for Speech Corpus Transcription Case of Arabic Algerian Dialects

2019-09-01 · WS 2019 9 · Ilyes Zine, Mohamed Cherif Zeghad, Soumia Bougrine, Hadda Cherroun

Towards a Corpus of Violence Acts in Arabic Social Media

2016-05-01 · LREC 2016 5 · Ayman Alhelbawy, Poesio Massimo, Udo Kruschwitz

In this paper we present a new corpus of Arabic tweets that mention some form of violent event, developed to support the automatic identification of Human Rights Abuse. The dataset was manually labelled for seven classes…

Form