paper-with-me

Papers

Compilation of an Arabic Children's Corpus

2016-05-01 · LREC 2016 5 · Latifa Al-Sulaiti, Noorhan Abbas, Claire Brierley, Eric Atwell, Ayman Alghamdi

Inspired by the Oxford Children{'}s Corpus, we have developed a prototype corpus of Arabic texts written and/or selected for children. Our Arabic Children{'}s Corpus of 2950 documents and nearly 2 million words has been collected manually from the web during a 3-month project. It is of high quality, and contains a range of different children{'}s genres based on sources located, including classic tales from The Arabian Nights, and popular fictional characters such as Goha. We anticipate that the current and subsequent versions of our corpus will lead to interesting studies in text classification, language use, and ideology in children{'}s texts.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

General Classificationtext-classificationText Classification

Similar Papers 제목 키워드 기반

The International Corpus of Arabic: Compilation, Analysis and Evaluation

2014-10-01 · WS 2014 10 · Sameh Alansary, Magdy Nagi
Morphological Analysis

A Game with a Purpose for Automatic Detection of Children's Speech Disabilities using Limited Speech Resources

2017-09-01 · RANLP 2017 9 · Reem Salem, Mohamed Elmahdy, Slim Abdennadher, Injy Hamed

Speech therapists and researchers are becoming more concerned with the use of computer-based systems in the therapy of speech disorders. In this paper, we propose a computer-based game with a purpose (GWAP) for speech th…

Information Retrieval

Building a Non-native Speech Corpus Featuring Chinese-English Bilingual Children: Compilation and Rationale

2023-04-30 · Hiuchung Hung, Andreas Maier, Thorsten Piske

This paper introduces a non-native speech corpus consisting of narratives from fifty 5- to 6-year-old Chinese-English children. Transcripts totaling 6.5 hours of children taking a narrative comprehension test in English …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Leveraging Domain Adaptation and Data Augmentation to Improve Qur'anic IR in English and Arabic

2023-12-05 · Vera Pavlova

In this work, we approach the problem of Qur'anic information retrieval (IR) in Arabic and English. Using the latest state-of-the-art methods in neural IR, we research what helps to tackle this task more efficiently. Tra…

Data AugmentationDomain AdaptationInformation RetrievalLanguage Modelling+1

Tharwa: A Large Scale Dialectal Arabic - Standard Arabic - English Lexicon

2014-05-01 · LREC 2014 5 · Mona Diab, Mohamed Al-Badrashiny, Maryam Aminian, Mohammed Attia 외

We introduce an electronic three-way lexicon, Tharwa, comprising Dialectal Arabic, Modern Standard Arabic and English correspondents. The paper focuses on Egyptian Arabic as the first pilot dialect for the resource, with…

POS