paper-with-me

Papers

When Scale Meets Diversity: Evaluating Language Models on Fine-Grained Multilingual Claim Verification

2025-07-28 · Hanna Shcharbakova, Tatiana Anikina, Natalia Skachkova, Josef van Genabith arxiv

The rapid spread of multilingual misinformation requires robust automated fact verification systems capable of handling fine-grained veracity assessments across diverse languages. While large language models have shown remarkable capabilities across many NLP tasks, their effectiveness for multilingual claim verification with nuanced classification schemes remains understudied. We conduct a comprehensive evaluation of five state-of-the-art language models on the X-Fact dataset, which spans 25 languages with seven distinct veracity categories. Our experiments compare small language models (encoder-based XLM-R and mT5) with recent decoder-only LLMs (Llama 3.1, Qwen 2.5, Mistral Nemo) using both prompting and fine-tuning approaches. Surprisingly, we find that XLM-R (270M parameters) substantially outperforms all tested LLMs (7-12B parameters), achieving 57.7% macro-F1 compared to the best LLM performance of 16.9%. This represents a 15.8% improvement over the previous state-of-the-art (41.9%), establishing new performance benchmarks for multilingual fact verification. Our analysis reveals problematic patterns in LLM behavior, including systematic difficulties in leveraging evidence and pronounced biases toward frequent categories in imbalanced data settings. These findings suggest that for fine-grained multilingual fact verification, smaller specialized models may be more effective than general-purpose large models, with important implications for practical deployment of fact-checking systems.

📄 PDF Abstract BibTeX arXiv:2507.20700

Code (0)

등록된 구현이 없습니다.

Tasks

Fact Verification

Similar Papers 제목 키워드 기반

When Large Vision-Language Model Meets Large Remote Sensing Imagery: Coarse-to-Fine Text-Guided Token Pruning

2025-03-10 · Junwei Luo, Yingying Zhang, Xue Yang, Kang Wu 외

Efficient vision-language understanding of large Remote Sensing Images (RSIs) is meaningful but challenging. Current Large Vision-Language Models (LVLMs) typically employ limited pre-defined grids to process images, lead…

Language ModelingLanguage ModellingToken ReductionVisual Question Answering (VQA)

When Molecular GAN Meets Byte-Pair Encoding

2024-09-29 · Huidong Tang, Chen Li, Yasuhiko Morimoto

Deep generative models, such as generative adversarial networks (GANs), are pivotal in discovering novel drug-like candidates via de novo molecular generation. However, traditional character-wise tokenizers often struggl…

Computational EfficiencyDiversity

The Tower of Babel Meets Web 2.0: User-Generated Content and its Applications in a Multilingual Context

2019-04-02 · B. Hecht, D. Gergle

This study explores language's fragmenting effect on user-generated content by examining the diversity of knowledge representations across 25 different Wikipedia language editions. This diversity is measured at two level…

DiversityWorld Knowledge

EfficientOCR: An Extensible, Open-Source Package for Efficiently Digitizing World Knowledge

2023-10-16 · Tom Bryan, Jacob Carlson, Abhishek Arora, Melissa Dell

Billions of public domain documents remain trapped in hard copy or lack an accurate digitization. Modern natural language processing methods cannot be used to index, retrieve, and summarize their texts; conduct computati…

Image RetrievalLanguage ModelingLanguage ModellingOptical Character Recognition+2

MHEntropy: Entropy Meets Multiple Hypotheses for Pose and Shape Recovery

2023-01-01 · ICCV 2023 1 · Rongyu Chen, Linlin Yang, Angela Yao

For monocular RGB-based 3D pose and shape estimation, multiple solutions are often feasible due to factors like occlusion and truncation. This work presents a multi-hypothesis probabilistic framework by optimizing th…

3D Hand Pose Estimation3D Human Pose Estimation3D Human ReconstructionDiversity+2