MegaWika 2: A More Comprehensive Multilingual Collection of Articles and their Sources
We introduce MegaWika 2, a large, multilingual dataset of Wikipedia articles with their citations and scraped web sources; articles are represented in a rich data structure, and scraped source texts are stored inline with precise character offsets of their citations in the article text. MegaWika 2 is a major upgrade from the original MegaWika, spanning six times as many articles and twice as many fully scraped citations. Both MegaWika and MegaWika 2 support report generation research ; whereas MegaWika also focused on supporting question answering and retrieval applications, MegaWika 2 is designed to support fact checking and analyses across time and language.
Code (0)
등록된 구현이 없습니다.
Tasks
Question AnsweringFact CheckingSimilar Papers 제목 키워드 기반
MegaWika: Millions of reports and their sources across 50 diverse languages
To foster the development of new models for collaborative AI-assisted report generation, we introduce MegaWika, consisting of 13 million Wikipedia articles in 50 diverse languages, along with their 71 million referenced …
ArticlesCross-Lingual Question AnsweringQuestion AnsweringRetrieval+1MultiNews: A Web collection of an Aligned Multimodal and Multilingual Corpus
Integrating Natural Language Processing (NLP) and computer vision is a promising effort. However, the applicability of these methods directly depends on the availability of a specific multimodal data that includes images…
ArticlesContent-Based Image RetrievalImage RetrievalMachine Translation+1A tool for enhanced search of multilingual digital libraries of e-journals
This paper outlines the main features of Bibli{\v{s}}a, a tool that offers various possibilities of enhancing queries submitted to large collections of TMX documents generated from aligned parallel articles residing in m…
ArticlesInformation RetrievalMachine TranslationNews Across Languages - Cross-Lingual Document Similarity and Event Tracking
In today's world, we follow news which is distributed globally. Significant events are reported by different sources and in different languages. In this work, we address the problem of tracking of events in a large multi…
ArticlesMultilingual, Multimodal Pipeline for Creating Authentic and Structured Fact-Checked Claim Dataset
The rapid proliferation of misinformation across online platforms underscores the urgent need for robust, up-to-date, explainable, and multilingual fact-checking resources. However, existing datasets are limited in scope…