paper-with-me

홈 › Papers

MegaWika 2: A More Comprehensive Multilingual Collection of Articles and their Sources

2025-08-05 · Samuel Barham, Chandler May, Benjamin Van Durme arxiv

We introduce MegaWika 2, a large, multilingual dataset of Wikipedia articles with their citations and scraped web sources; articles are represented in a rich data structure, and scraped source texts are stored inline with precise character offsets of their citations in the article text. MegaWika 2 is a major upgrade from the original MegaWika, spanning six times as many articles and twice as many fully scraped citations. Both MegaWika and MegaWika 2 support report generation research ; whereas MegaWika also focused on supporting question answering and retrieval applications, MegaWika 2 is designed to support fact checking and analyses across time and language.

📄 PDF Abstract BibTeX arXiv:2508.03828

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringFact Checking

Similar Papers 제목 키워드 기반

MegaWika: Millions of reports and their sources across 50 diverse languages

2023-07-13 · Samuel Barham, Orion Weller, Michelle Yuan, Kenton Murray 외

To foster the development of new models for collaborative AI-assisted report generation, we introduce MegaWika, consisting of 13 million Wikipedia articles in 50 diverse languages, along with their 71 million referenced …

ArticlesCross-Lingual Question AnsweringQuestion AnsweringRetrieval+1

MultiNews: A Web collection of an Aligned Multimodal and Multilingual Corpus

2017-11-01 · WS 2017 11 · Haithem Afli, Pintu Lohar, Andy Way

Integrating Natural Language Processing (NLP) and computer vision is a promising effort. However, the applicability of these methods directly depends on the availability of a specific multimodal data that includes images…

ArticlesContent-Based Image RetrievalImage RetrievalMachine Translation+1

A tool for enhanced search of multilingual digital libraries of e-journals

2012-05-01 · LREC 2012 5 · Ranka Stankovi{\'c}, Cvetana Krstev, Ivan Obradovi{\'c}, Aleks Trtovac 외

This paper outlines the main features of Bibli{\v{s}}a, a tool that offers various possibilities of enhancing queries submitted to large collections of TMX documents generated from aligned parallel articles residing in m…

ArticlesInformation RetrievalMachine Translation

News Across Languages - Cross-Lingual Document Similarity and Event Tracking

2015-12-22 · Jan Rupnik, Andrej Muhic, Gregor Leban, Primoz Skraba 외

In today's world, we follow news which is distributed globally. Significant events are reported by different sources and in different languages. In this work, we address the problem of tracking of events in a large multi…

Articles

Multilingual, Multimodal Pipeline for Creating Authentic and Structured Fact-Checked Claim Dataset

2026-01-12 · Z. Melce Hüsünbeyi, Virginie Mouilleron, Leonie Uhling, Daniel Foppe 외 arxiv

The rapid proliferation of misinformation across online platforms underscores the urgent need for robust, up-to-date, explainable, and multilingual fact-checking resources. However, existing datasets are limited in scope…