paper-with-me

홈 › Papers

Think Through Uncertainty: Improving Long-Form Generation Factuality via Reasoning Calibration

2026-04-13 · Xin Liu, Lu Wang arxiv

Large language models (LLMs) often hallucinate in long-form generation. Existing approaches mainly improve factuality through post-hoc revision or reinforcement learning (RL) with correctness-based rewards, but they do not teach the model to estimate which parts of its generation are reliable. As a result, models may still state incorrect claims confidently in their responses. Recent advances in reasoning have significantly improved LLM performance, and have been leveraged to estimate confidence by incorporating calibration into RL objectives. However, existing approaches remain limited to a single scalar confidence for the entire response, which is insufficient for long-form generation where uncertainty varies across individual claims. To mitigate this problem, we propose CURE, a framework that improves long-form factuality by teaching LLMs to reason about uncertainty at the claim level. We first introduce a Claim-Aware Reasoning Protocol, which structures outputs into atomic claims paired with explicit confidence estimates. We then develop a multi-stage training pipeline that aligns model confidence with claims' correctness and then optimizes on factuality. The resulting calibrated confidence further enables selective prediction, allowing the model to abstain from uncertain claims at inference time. Experiments on four long-form factuality benchmarks show that CURE consistently improves factual accuracy over competitive supervised and RL baselines, while maintaining factual recall. In particular, it improves claim-level accuracy by up to 39.9% on Biography generation. These gains are accompanied by improved calibration, as reflected by a 16.0% increase in AUROC on FactBench.

📄 PDF Abstract BibTeX arXiv:2604.12046

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Thinking Fast and Laterally: Multi-Agentic Approach for Reasoning about Uncertain Emerging Events

2024-12-10 · Stefan Dernbach, Alejandro Michel, Khushbu Agarwal, Christopher Brissette 외

This paper introduces lateral thinking to implement System-2 reasoning capabilities in AI systems, focusing on anticipatory and causal reasoning under uncertainty. We present a framework for systematic generation and mod…

ManagementSpecificity

ReQuest: Rethinking-based Question-Aware Frame Selection for Long-Form Video QA

2026-07-02 · Minkuk Kim, Suyong Yun, Young Tae Kim, Jinyoung Moon 외 arxiv

Recent multimodal large language models (MLLMs) have substantially advanced video understanding, yet long-form video QA remains challenging under fixed input token budgets, where uniform sampling can be inefficient for e…

ControlThinker: Unveiling Latent Semantics for Controllable Image Generation through Visual Reasoning

2025-06-04 · Feng Han, Yang Jiao, Shaoxiang Chen, Junhao Xu 외

The field of controllable image generation has seen significant advancements, with various architectures improving generation layout consistency with control signals. However, contemporary methods still face challenges i…

Image GenerationVisual Reasoning

When to Think and When to Look: Uncertainty-Guided Lookback

2025-11-19 · Jing Bi, Filippos Bellos, Junjia Guo, Yayuan Li 외 arxiv

Test-time thinking (that is, generating explicit intermediate reasoning chains) is known to boost performance in large language models and has recently shown strong gains for large vision language models (LVLMs). However…

Visual GroundingVisual Reasoning

OmniThink: Expanding Knowledge Boundaries in Machine Writing through Thinking

2025-01-16 · Zekun Xi, Wenbiao Yin, Jizhan Fang, Jialong Wu 외

Machine writing with large language models often relies on retrieval-augmented generation. However, these approaches remain confined within the boundaries of the model's predefined scope, limiting the generation of conte…

ArticlesRetrieval-augmented Generation