Y-NQ: English-Yorùbá Evaluation dataset for Open-Book Reading Comprehension and Text Generation
The purpose of this work is to share an English-Yor\ub\'a evaluation dataset for open-book reading comprehension and text generation to assess the performance of models both in a high- and a low- resource language. The dataset contains 358 questions and answers on 338 English documents and 208 Yor\ub\'a documents. The average document length is ~ 10k words for English and 430 words for Yor\ub\'a. Experiments show a consistent disparity in performance between the two languages, with Yor\ub\'a falling behind English for automatic metrics even if documents are much shorter for this language. For a small set of documents with comparable length, performance of Yor\ub\'a drops by x2.5 times. When analyzing performance by length, we observe that Yor\ub\'a decreases performance dramatically for documents that reach 1500 words while English performance is barely affected at that length. Our dataset opens the door to showcasing if English LLM reading comprehension capabilities extend to Yor\`ub\'a, which for the evaluated LLMs is not the case.
Code (0)
등록된 구현이 없습니다.
Tasks
Reading ComprehensionText GenerationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
OpenStaxQA: A multilingual dataset based on open-source college textbooks
We present OpenStaxQA, an evaluation benchmark specific to college-level educational applications based on 43 open-source college textbooks in English, Spanish, and Polish, available under a permissive Creative Commons l…
Detection of Reading Absorption in User-Generated Book Reviews: Resources Creation and Evaluation
To detect how and when readers are experiencing engagement with a literary work, we bring together empirical literary studies and language technology via focusing on the affective state of absorption. The goal of our res…
Binary ClassificationSentenceSentence EmbeddingSentence-EmbeddingProsody Analysis of Audiobooks
Recent advances in text-to-speech have made it possible to generate natural-sounding audio from text. However, audiobook narrations involve dramatic vocalizations and intonations by the reader, with greater reliance on e…
AttributeLanguage ModelingLanguage ModellingProsody Prediction+2AI translation of literary texts is "fine", but readers still prefer human translations
AI translation of literary works is increasingly common. While the content may be rendered adequately, we do not know enough about how readers experience it in terms of immersiveness and literary effect-aspects poorly ca…
Machine TranslationCalliope: A TTS-based Narrated E-book Creator Ensuring Exact Synchronization, Privacy, and Layout Fidelity
A narrated e-book combines synchronized audio with digital text, highlighting the currently spoken word or sentence during playback. This format supports early literacy and assists individuals with reading challenges, wh…