paper-with-me

Papers

Leveraging LLM For Synchronizing Information Across Multilingual Tables

2025-04-03 · Siddharth Khincha, Tushar Kataria, Ankita Anand, Dan Roth, Vivek Gupta

The vast amount of online information today poses challenges for non-English speakers, as much of it is concentrated in high-resource languages such as English and French. Wikipedia reflects this imbalance, with content in low-resource languages frequently outdated or incomplete. Recent research has sought to improve cross-language synchronization of Wikipedia tables using rule-based methods. These approaches can be effective, but they struggle with complexity and generalization. This paper explores large language models (LLMs) for multilingual information synchronization, using zero-shot prompting as a scalable solution. We introduce the Information Updation dataset, simulating the real-world process of updating outdated Wikipedia tables, and evaluate LLM performance. Our findings reveal that single-prompt approaches often produce suboptimal results, prompting us to introduce a task decomposition strategy that enhances coherence and accuracy. Our proposed method outperforms existing baselines, particularly in Information Updation (1.79%) and Information Addition (20.58%), highlighting the model strength in dynamically updating and enriching data across architectures

📄 PDF Abstract BibTeX arXiv:2504.02559

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

InfoSync: Information Synchronization across Multilingual Semi-structured Tables

2023-07-06 · Siddharth Khincha, Chelsi Jain, Vivek Gupta, Tushar Kataria 외

Information Synchronization of semi-structured data across languages is challenging. For instance, Wikipedia tables in one language should be synchronized across languages. To address this problem, we introduce a new dat…

Factual Inconsistencies in Multilingual Wikipedia Tables

2025-07-24 · Silvia Cappa, Lingxiao Kong, Pille-Riin Peet, Fanfu Wei 외 arxiv

Wikipedia serves as a globally accessible knowledge source with content in over 300 languages. Despite covering the same topics, the different versions of Wikipedia are written and updated independently. This leads to fa…

PulseBench-Tab: A Multilingual Benchmark for Table Extraction with Graph-Based Evaluation

2026-04-21 · Ritvik Pandey, Sid Manchkanti, Mohammed Wazir Adain, Mohammed Hadi 외 arxiv

We introduce PulseBench-Tab, an open multilingual benchmark for evaluating table extraction from document images. The benchmark comprises 1,820 human-annotated tables spanning 9 languages and 4 scripts (Latin, CJK, Arabi…

IndicDB -- Benchmarking Multilingual Text-to-SQL Capabilities in Indian Languages

2026-04-15 · Aviral Dawar, Roshan Karanth, Vikram Goyal, Dhruv Kumar arxiv

While Large Language Models (LLMs) have significantly advanced Text-to-SQL performance, existing benchmarks predominantly focus on Western contexts and simplified schemas, leaving a gap in real-world, non-Western applica…

Semantic Parsing

Creating Large-Scale Multilingual Cognate Tables

2018-05-01 · LREC 2018 5 · Winston Wu, David Yarowsky
Machine TranslationSemantic Textual SimilarityTransliteration