paper-with-me

홈 › Papers

AEScorer: An Agentic Evidence-Grounded Framework for Graded Factuality Verification

2026-01-07 · Hui Huang, Muyun Yang, Yuki Arase arxiv

Despite the significant advancements of Large Language Models (LLMs), their factuality remains a critical challenge, creating a growing need for more nuanced factuality verification. Existing factuality verification methods do not capture graded judgments, even though factuality is better understood as a spectrum rather than a binary of right and wrong. To bridge this gap, we focus on graded factuality verification and propose AEScorer, an agentic evidence-grounded framework with two stages: agentic evidence acquisition and graded scoring. AEScorer first gathers and refines external evidence through agentic search, and then predicts a scalar factuality score to distinguish nuanced differences in factual correctness. We further construct GradedVeriBench, a benchmark for graded factuality verification spanning both general and multi-hop question answering. Experimental results on GradedVeriBench show that AEScorer substantially outperforms existing methods across both settings, demonstrating the value of coupling targeted evidence acquisition with graded scoring.

📄 PDF Abstract BibTeX arXiv:2601.03605

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Multimodal Agentic Pathology Co-pilot via Evidence Grounded Reasoning

2026-06-06 · Zhe Xu, Zhengyu Zhang, Zhiyuan Cai, Jiahao Xu 외 arxiv

Pathology is the cornerstone of modern medicine, where accurate decision-making relies heavily on evidence-based practices. While artificial intelligence (AI) has the potential to transform clinical workflows, the inters…

Agentic Hybrid RAG for Evidence-Grounded Muon Collider Analysis

2026-06-09 · Ruobing Jiang, Dawei Fu, Cheng Jiang, Tianyi Yang 외 arxiv

Muon collider research spans accelerator physics, detector instrumentation, and high-energy phenomenology, with relevant evidence scattered across a rapidly expanding and heterogeneous body of scientific literature. As h…

Semantic RetrievalQuestion AnsweringAnswer Generation

CARE: Towards Clinical Accountability in Multi-Modal Medical Reasoning with an Evidence-Grounded Agentic Framework

2026-03-02 · Yuexi Du, Jinglu Wang, Shujie Liu, Nicha C. Dvornek 외 arxiv

Large visual language models (VLMs) have shown strong multi-modal medical reasoning ability, but most operate as end-to-end black boxes, diverging from clinicians' evidence-based, staged workflows and hindering clinical …

Reinforcement LearningVisual Grounding

VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos

2026-02-08 · Wenqi Liu, Yunxiao Wang, Shijie Ma, Meng Liu 외 arxiv

In long-video understanding, conventional uniform frame sampling often fails to capture key visual evidence, leading to degraded performance and increased hallucinations. To address this, recent agentic thinking-with-vid…

Reinforcement LearningQuestion AnsweringVideo Grounding

BALMS: Benchmarking Agentic LLMs for Longitudinal Mental Health Sensing

2026-08-27 · Yu Yvonne Wu, Arvind Pillai, Yuliang Chen, Yuwei Zhang 외 arxiv

Mental health assessment relies on episodic self-report scales, which convert subjective states such as stress into numerical scores but provide only sparse snapshots of wellbeing. Wearable devices offer longitudinal beh…

Natural Language Queries