paper-with-me

홈 › Papers

An Empirical Analysis of Factual Errors in Human-Written Text and its Application

2026-06-26 · Kazuma Iwamoto, Kazumasa Omura, Shotaro Ishihara arxiv

Factual Error Detection (FED), which is the task of identifying factually incorrect spans in a given text, has long been recognized as an important research problem. However, with the rapid rise of large language models (LLMs), research attention has shifted toward factual errors specific to LLM-generated text (hallucinations) and their detection. As a result, the detection of factual errors in human-written text has been relatively neglected. To address this gap, we first distill a taxonomy of human-induced factual errors by analyzing corrections of newspaper articles, a representative source of text that is guaranteed to be human-written and contains few grammatical errors. Our analysis revealed that there are characteristic categories such as kanji misconversions and numeral classifier errors, which are not focused in existing hallucination benchmarks. Based on the taxonomy, we then evaluate the FED capability of vanilla LLMs on synthesized realistic test cases and real corrections. Experimental results demonstrated that even high-performance LLMs such as GPT-5.4 achieved only word-level F1 score of 52% on the synthetic evaluation data, highlighting the task difficulty. Furthermore, a detailed analysis by detection difficulty revealed the current state of FED.

📄 PDF Abstract BibTeX arXiv:2606.27959

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AI vs. Human -- Differentiation Analysis of Scientific Content Generation

2023-01-24 · Yongqiang Ma, Jiawei Liu, Fan Yi, Qikai Cheng 외

Recent neural language models have taken a significant step forward in producing remarkably controllable, fluent, and grammatical text. Although studies have found that AI-generated text is not distinguishable from human…

Text Detection

FLEEK: Factual Error Detection and Correction with Evidence Retrieved from External Knowledge

2023-10-26 · Farima Fatahi Bayat, Kun Qian, Benjamin Han, Yisi Sang 외

Detecting factual errors in textual information, whether generated by large language models (LLM) or curated by humans, is crucial for making informed decisions. LLMs' inability to attribute their claims to external know…

Attribute

Annotating Errors in English Learners' Written Language Production: Advancing Automated Written Feedback Systems

2025-08-09 · Steven Coyne, Diana Galvan-Sosa, Ryan Spring, Camélia Guerraoui 외 arxiv

Recent advances in natural language processing (NLP) have contributed to the development of automated writing evaluation (AWE) systems that can correct grammatical errors. However, while these systems are effective at im…

Measuring and Reducing LLM Hallucination without Gold-Standard Answers

2024-02-16 · Jiaheng Wei, Yuanshun Yao, Jean-Francois Ton, Hongyi Guo 외

LLM hallucination, i.e. generating factually incorrect yet seemingly convincing answers, is currently a major threat to the trustworthiness and reliability of LLMs. The first step towards solving this complicated problem…

HallucinationIn-Context Learning

Neural Deepfake Detection with Factual Structure of Text

2020-10-15 · EMNLP 2020 11 · Wanjun Zhong, Duyu Tang, Zenan Xu, Ruize Wang 외

Deepfake detection, the task of automatically discriminating machine-generated text, is increasingly critical with recent advances in natural language generative models. Existing approaches to deepfake detection typicall…

DeepFake DetectionFace SwappingGraph Neural NetworkSentence