paper-with-me

홈 › Papers

Comparing Hallucination Detection Metrics for Multilingual Generation

2024-02-16 · Haoqiang Kang, Terra Blevins, Luke Zettlemoyer

While many hallucination detection techniques have been evaluated on English text, their effectiveness in multilingual contexts remains unknown. This paper assesses how well various factual hallucination detection metrics (lexical metrics like ROUGE and Named Entity Overlap, and Natural Language Inference (NLI)-based metrics) identify hallucinations in generated biographical summaries across languages. We compare how well automatic metrics correlate to each other and whether they agree with human judgments of factuality. Our analysis reveals that while the lexical metrics are ineffective, NLI-based metrics perform well, correlating with human annotations in many settings and often outperforming supervised models. However, NLI metrics are still limited, as they do not detect single-fact hallucinations well and fail for lower-resource languages. Therefore, our findings highlight the gaps in exisiting hallucination detection methods for non-English languages and motivate future research to develop more robust multilingual detection methods for LLM hallucinations.

📄 PDF Abstract BibTeX arXiv:2402.10496

Code (0)

등록된 구현이 없습니다.

Tasks

HallucinationNatural Language InferenceSentence

Similar Papers 제목 키워드 기반

Multilingual Hallucination Gaps in Large Language Models

2024-10-23 · Cléa Chataigner, Afaf Taïk, Golnoosh Farnadi

Large language models (LLMs) are increasingly used as alternatives to traditional search engines given their capacity to generate text that resembles human language. However, this shift is concerning, as LLMs often gener…

HallucinationText Generation

BHRAM-IL: A Benchmark for Hallucination Recognition and Assessment in Multiple Indian Languages

2025-12-01 · Hrishikesh Terdalkar, Kirtan Bhojani, Aryan Dongare, Omm Aditya Behera arxiv

Large language models (LLMs) are increasingly deployed in multilingual applications but often generate plausible yet incorrect or misleading outputs, known as hallucinations. While hallucination detection has been studie…

When Models Lie, We Learn: Multilingual Span-Level Hallucination Detection with PsiloQA

2025-10-06 · Elisei Rykov, Kseniia Petrushina, Maksim Savkin, Valerii Olisov 외 arxiv

Hallucination detection remains a fundamental challenge for the safe and reliable deployment of large language models (LLMs), especially in applications requiring factual accuracy. Existing hallucination benchmarks often…

Multilingual Fine-Grained News Headline Hallucination Detection

2024-07-22 · Jiaming Shen, Tianqi Liu, Jialu Liu, Zhen Qin 외

The popularity of automated news headline generation has surged with advancements in pre-trained language models. However, these models often suffer from the ``hallucination'' problem, where the generated headline is not…

HallucinationHeadline GenerationIn-Context Learning

Halluverse-M^3: A multitask multilingual benchmark for hallucination in LLMs

2026-02-06 · Samir Abdaljalil, Parichit Sharma, Erchin Serpedin, Hasan Kurban arxiv

Hallucinations in large language models remain a persistent challenge, particularly in multilingual and generative settings where factual consistency is difficult to maintain. While recent models show strong performance …

Question Answering