paper-with-me

Papers

Factual Inconsistencies in Multilingual Wikipedia Tables

2025-07-24 · Silvia Cappa, Lingxiao Kong, Pille-Riin Peet, Fanfu Wei, Yuchen Zhou, Jan-Christoph Kalo arxiv

Wikipedia serves as a globally accessible knowledge source with content in over 300 languages. Despite covering the same topics, the different versions of Wikipedia are written and updated independently. This leads to factual inconsistencies that can impact the neutrality and reliability of the encyclopedia and AI systems, which often rely on Wikipedia as a main training source. This study investigates cross-lingual inconsistencies in Wikipedia's structured content, with a focus on tabular data. We developed a methodology to collect, align, and analyze tables from Wikipedia multilingual articles, defining categories of inconsistency. We apply various quantitative and qualitative metrics to assess multilingual alignment using a sample dataset. These insights have implications for factual verification, multilingual knowledge interaction, and design for reliable AI systems leveraging Wikipedia content.

📄 PDF Abstract BibTeX arXiv:2507.18406

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

InfoSync: Information Synchronization across Multilingual Semi-structured Tables

2023-07-06 · Siddharth Khincha, Chelsi Jain, Vivek Gupta, Tushar Kataria 외

Information Synchronization of semi-structured data across languages is challenging. For instance, Wikipedia tables in one language should be synchronized across languages. To address this problem, we introduce a new dat…

Zoom In Disparities in Healthcare LLM Q&A

2025-10-20 · Ipek Baris Schlicht, Burcu Sayin, Zhixue Zhao, Frederik M. Labonté 외 arxiv

Equitable access to reliable health information is vital when integrating AI into healthcare. Yet, information quality varies across languages, raising concerns about the reliability and consistency of multilingual Large…

Detecting Corpus-Level Knowledge Inconsistencies in Wikipedia with Large Language Models

2025-09-27 · Sina J. Semnani, Jirayu Burapacheep, Arpandeep Khatua, Thanawan Atchariyachanvanit 외 arxiv

Wikipedia is the largest open knowledge corpus, widely used worldwide and serving as a key resource for training large language models (LLMs) and retrieval-augmented generation (RAG) systems. Ensuring its accuracy is the…

Leveraging LLM For Synchronizing Information Across Multilingual Tables

2025-04-03 · Siddharth Khincha, Tushar Kataria, Ankita Anand, Dan Roth 외

The vast amount of online information today poses challenges for non-English speakers, as much of it is concentrated in high-resource languages such as English and French. Wikipedia reflects this imbalance, with content …

Detecting Knowledge Inconsistencies Across Text, Tables, and Knowledge Graphs

2026-07-28 · Fanfu Wei, Thibault Ehrhart, Raphaël Troncy arxiv

Wikipedia and Wikidata are widely used for information access, LLM pre-training, and retrieval-augmented generation. Their knowledge is deeply connected but scattered across text, tables, and knowledge graphs. This raise…

Knowledge Graphs