Project PIAF: Building a Native French Question-Answering Dataset
Motivated by the lack of data for non-English languages, in particular for the evaluation of downstream tasks such as Question Answering, we present a participatory effort to collect a native French Question Answering Dataset. Furthermore, we describe and publicly release the annotation tool developed for our collection effort, along with the data obtained and preliminary baselines.
Code (1)
Tasks
Question AnsweringSimilar Papers 제목 키워드 기반
PIAFusion: A progressive infrared and visible image fusion network based on illumination aware
Infrared and visible image fusion aims to synthesize a single fused image containing salient targets and abundant texture details even under extreme illumination conditions. However, existing image fusion algorithms fail…
Infrared And Visible Image FusionSemantic SegmentationBuilding a Bilingual Vietnamese-French Named Entity Annotated Corpus through Cross-Linguistic Projection
The creation of high-quality named entity annotated resources is time-consuming and an expensive process. Most of the gold standard corpora are available for English but not for less-resourced languages such as Vietnames…
Hard Time Parsing Questions: Building a QuestionBank for French
We present the French Question Bank, a treebank of 2600 questions. We show that classical parsing model performance drop while the inclusion of this data set is highly beneficial without harming the parsing of non-questi…
Synthetic Data Augmentation for Zero-Shot Cross-Lingual Question Answering
Coupled with the availability of large scale datasets, deep learning architectures have enabled rapid progress on the Question Answering task. However, most of those datasets are in English, and the performances of state…
Cross-Lingual Question AnsweringData AugmentationQuestion AnsweringQuestion Generation+1An Aligned French-Chinese corpus of 10K segments from university educational material
This paper describes a corpus of nearly 10K French-Chinese aligned segments, produced by post-editing machine translated computer science courseware. This corpus was built from 2013 to 2016 within the PROJECT{\_}NAME pro…
Machine TranslationTranslation