paper-with-me

Papers

Multi-round, Chain-of-thought Post-editing for Unfaithful Summaries

2025-01-20 · Yi-Hui Lee, Xiangci Li, Jessica Ouyang

Recent large language models (LLMs) have demonstrated a remarkable ability to perform natural language understanding and generation tasks. In this work, we investigate the use of LLMs for evaluating faithfulness in news summarization, finding that it achieves a strong correlation with human judgments. We further investigate LLMs' capabilities as a faithfulness post-editor, experimenting with different chain-of-thought prompts for locating and correcting factual inconsistencies between a generated summary and the source news document and are able to achieve a higher editing success rate than was reported in prior work. We perform both automated and human evaluations of the post-edited summaries, finding that prompting LLMs using chain-of-thought reasoning about factual error types is an effective faithfulness post-editing strategy, performing comparably to fine-tuned post-editing models. We also demonstrate that multiple rounds of post-editing, which has not previously been explored, can be used to gradually improve the faithfulness of summaries whose errors cannot be fully corrected in a single round.

📄 PDF Abstract BibTeX arXiv:2501.11273

Code (0)

등록된 구현이 없습니다.

Tasks

Natural Language UnderstandingNews Summarization

Similar Papers 제목 키워드 기반

Instruction-based Image Editing with Planning, Reasoning, and Generation

2026-02-26 · Liya Ji, Chenyang Qi, Qifeng Chen arxiv

Editing images via instruction provides a natural way to generate interactive content, but it is a big challenge due to the higher requirement of scene understanding and generation. Prior work utilizes a chain of large l…

Object SegmentationScene UnderstandingImage Editing

Verify-and-Edit: A Knowledge-Enhanced Chain-of-Thought Framework

2023-05-05 · Ruochen Zhao, Xingxuan Li, Shafiq Joty, Chengwei Qin 외

As large language models (LLMs) have become the norm in NLP, demonstrating good performance in generation and reasoning tasks, one of its most fatal disadvantages is the lack of factual correctness. Generating unfactual …

Open-Domain Question AnsweringQuestion Answering

RIPPLECOT: Amplifying Ripple Effect of Knowledge Editing in Language Models via Chain-of-Thought In-Context Learning

2024-10-04 · Zihao Zhao, Yuchen Yang, Yijiang Li, Yinzhi Cao

The ripple effect poses a significant challenge in knowledge editing for large language models. Namely, when a single fact is edited, the model struggles to accurately update the related facts in a sequence, which is eva…

In-Context Learningknowledge editing

CoRGI: Verified Chain-of-Thought Reasoning with Post-hoc Visual Grounding

2025-08-01 · Shixin Yi, Lin Shang arxiv

Multimodal reasoning with vision-language models (VLMs) often suffers from hallucinations, as models tend to generate explanations after only a superficial inspection of the image. We present \textbf{CoRGI}(\textbf{C}hai…

Multimodal ReasoningVisual Grounding

Universal Image Restoration via Internalized Chain-of-Thought Reasoning

2026-06-16 · Yu Guo, Zhengru Fang, Shengfeng He, Senkang Hu 외 arxiv

Image restoration seeks to recover high-quality images from degraded inputs but becomes highly ill-posed under complex, mixed degradations. While unified all-in-one models are common, their performance declines as degrad…

Image RestorationImage Editing