paper-with-me

홈 › Papers

When Transliteration Met Crowdsourcing : An Empirical Study of Transliteration via Crowdsourcing using Efficient, Non-redundant and Fair Quality Control

2014-05-01 · LREC 2014 5 · Mitesh M. Khapra, Ananthakrishnan Ramanathan, Anoop Kunchukuttan, Karthik Visweswariah, Pushpak Bhattacharyya

Sufficient parallel transliteration pairs are needed for training state of the art transliteration engines. Given the cost involved, it is often infeasible to collect such data using experts. Crowdsourcing could be a cheaper alternative, provided that a good quality control (QC) mechanism can be devised for this task. Most QC mechanisms employed in crowdsourcing are aggressive (unfair to workers) and expensive (unfair to requesters). In contrast, we propose a low-cost QC mechanism which is fair to both workers and requesters. At the heart of our approach, lies a rule based Transliteration Equivalence approach which takes as input a list of vowels in the two languages and a mapping of the consonants in the two languages. We empirically show that our approach outperforms other popular QC mechanisms ({\textbackslash}textit{viz.}, consensus and sampling) on two vital parameters : (i) fairness to requesters (lower cost per correct transliteration) and (ii) fairness to workers (lower rate of rejecting correct answers). Further, as an extrinsic evaluation we use the standard NEWS 2010 test set and show that such quality controlled crowdsourced data compares well to expert data when used for training a transliteration engine.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

FairnessTransliteration

Similar Papers 제목 키워드 기반

A Deep Learning Based Approach to Transliteration

2018-07-01 · WS 2018 7 · Soumyadeep Kundu, Sayantan Paul, Santanu Pal

In this paper, we propose different architectures for language independent machine transliteration which is extremely important for natural language processing (NLP) applications. Though a number of statistical models fo…

Deep LearningInformation RetrievalMachine TranslationNMT+2

Design Challenges in Named Entity Transliteration

2018-08-07 · COLING 2018 8 · Yuval Merhav, Stephen Ash

We analyze some of the fundamental design challenges that impact the development of a multilingual state-of-the-art named entity transliteration system, including curating bilingual named entity datasets and evaluation o…

DecoderTransliteration

Hybrid approach for transliteration of Algerian arabizi: a primary study

2018-08-10 · Imane Guellil, Faical Azouaou, Fodil Benali, Ala-eddine Hachani 외

A hybrid approach for the transliteration of Algerian Arabizi: A primary study In this paper, we present a hybrid approach for the transliteration of the Algerian Arabizi. We define a set of rules enable us the passage f…

Transliteration

Statistical Models for Unsupervised, Semi-Supervised Supervised Transliteration Mining

2017-06-01 · CL 2017 6 · Hassan Sajjad, Helmut Schmid, Alex Fraser, er 외

We present a generative model that efficiently mines transliteration pairs in a consistent fashion in three different settings: unsupervised, semi-supervised, and supervised transliteration mining. The model interpolates…

Transliteration

Does Transliteration Help Multilingual Language Modeling?

2022-01-29 · Ibraheem Muhammad Moosa, Mahmud Elahi Akhter, Ashfia Binte Habib

Script diversity presents a challenge to Multilingual Language Models (MLLM) by reducing lexical overlap among closely related languages. Therefore, transliterating closely related languages that use different writing sc…

DiversityLanguage ModelingLanguage ModellingMultiple Choice Question Answering (MCQA)+5