paper-with-me

홈 › Papers

Code-Mixed Question Answering Challenge: Crowd-sourcing Data and Techniques

2018-07-01 · WS 2018 7 · Ch, Khyathi u, Ekaterina Loginova, Vishal Gupta, Josef van Genabith, G{\"u}nter Neumann, Manoj Chinnakotla, Eric Nyberg, Alan W. black

Code-Mixing (CM) is the phenomenon of alternating between two or more languages which is prevalent in bi- and multi-lingual communities. Most NLP applications today are still designed with the assumption of a single interaction language and are most likely to break given a CM utterance with multiple languages mixed at a morphological, phrase or sentence level. For example, popular commercial search engines do not yet fully understand the intents expressed in CM queries. As a first step towards fostering research which supports CM in NLP applications, we systematically crowd-sourced and curated an evaluation dataset for factoid question answering in three CM languages - Hinglish (Hindi+English), Tenglish (Telugu+English) and Tamlish (Tamil+English) which belong to two language families (Indo-Aryan and Dravidian). We share the details of our data collection process, techniques which were used to avoid inducing lexical bias amongst the crowd workers and other CM specific linguistic properties of the dataset. Our final dataset, which is available freely for research purposes, has 1,694 Hinglish, 2,848 Tamlish and 1,391 Tenglish factoid questions and their answers. We discuss the techniques used by the participants for the first edition of this ongoing challenge.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringSentence

Similar Papers 제목 키워드 기반

Uncovering Code-Mixed Challenges: A Framework for Linguistically Driven Question Generation and Neural Based Question Answering

2018-10-01 · CONLL 2018 10 · Deepak Gupta, Pabitra Lenka, Asif Ekbal, Pushpak Bhattacharyya

Existing research on question answering (QA) and comprehension reading (RC) are mainly focused on the resource-rich language like English. In recent times, the rapid growth of multi-lingual web content has posed several …

Question AnsweringQuestion GenerationQuestion-Generation

A Unified Framework for Multilingual and Code-Mixed Visual Question Answering

2020-12-01 · Asian Chapter of the Association for Computational Linguistics 2020 · Deepak Gupta, Pabitra Lenka, Asif Ekbal, Pushpak Bhattacharyya

In this paper, we propose an effective deep learning framework for multilingual and code- mixed visual question answering. The pro- posed model is capable of predicting answers from the questions in Hindi, English or Cod…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Evaluating Mixed-initiative Conversational Search Systems via User Simulation

2022-04-17 · Ivan Sekulić, Mohammad Aliannejadi, Fabio Crestani

Clarifying the underlying user information need by asking clarifying questions is an important feature of modern conversational search system. However, evaluation of such systems through answering prompted clarifying que…

Conversational SearchText GenerationUser Simulation

A Paraphrase Generation System for EHR Question Answering

2019-08-01 · WS 2019 8 · Sarvesh Soni, Kirk Roberts

This paper proposes a dataset and method for automatically generating paraphrases for clinical questions relating to patient-specific information in electronic health records (EHRs). Crowdsourcing is used to collect 10,5…

DecoderParaphrase GenerationQuestion Answering

SelQA: A New Benchmark for Selection-based Question Answering

2016-06-27 · Tomasz Jurczyk, Michael Zhai, Jinho D. Choi

This paper presents a new selection-based question answering dataset, SelQA. The dataset consists of questions generated through crowdsourcing and sentence length answers that are drawn from the ten most prevalent topics…

Question AnsweringSentence