paper-with-me

홈 › Papers

DocLens : A Tool-Augmented Multi-Agent Framework for Long Visual Document Understanding

2025-11-14 · Dawei Zhu, Rui Meng, Jiefeng Chen, Sujian Li, Tomas Pfister, Jinsung Yoon arxiv

Comprehending long visual documents, where information is distributed across extensive pages of text and visual elements, is a critical but challenging task for modern Vision-Language Models (VLMs). Existing approaches falter on a fundamental challenge: evidence localization. They struggle to retrieve relevant pages and overlook fine-grained details within visual elements, leading to limited performance and model hallucination. To address this, we propose DocLens, a tool-augmented multi-agent framework that effectively ``zooms in'' on evidence like a lens. It first navigates from the full document to specific visual elements on relevant pages, then employs a sampling-adjudication mechanism to generate a single, reliable answer. Paired with Gemini-2.5-Pro, DocLens achieves state-of-the-art performance on MMLongBench-Doc and FinRAGBench-V, surpassing even human experts. The framework's superiority is particularly evident on vision-centric and unanswerable queries, demonstrating the power of its enhanced localization capabilities.

📄 PDF Abstract BibTeX arXiv:2511.11552

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DocLens: Multi-aspect Fine-grained Evaluation for Medical Text Generation

2023-11-16 · Yiqing Xie, Sheng Zhang, Hao Cheng, PengFei Liu 외

Medical text generation aims to assist with administrative work and highlight salient information to support decision-making. To reflect the specific requirements of medical text, in this paper, we propose a set of metri…

Decision MakingInstruction FollowingText Generation

When Users Are Happy but Agents Are Wrong: Multi-Dimensional Evaluation of Tool-Augmented Dialogue

2025-10-22 · Tanya Shourya, Yingfan Wang, Zhaoyi Joey Hou, Shamik Roy 외 arxiv

Evaluating conversational AI systems that use external tools is challenging, as errors can arise from complex interactions among user, agent, and tools. While existing evaluation methods assess either user satisfaction o…

PDE-Agent: A toolchain-augmented multi-agent framework for PDE solving

2025-12-18 · Jianming Liu, Ren Zhu, Jian Xu, Kun Ding 외 arxiv

Solving Partial Differential Equations (PDEs) is a cornerstone of engineering and scientific research. Traditional methods for PDE solving are cumbersome, relying on manual setup and domain expertise. While Physics-Infor…

Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents

2025-10-03 · Wonjoong Kim, Sangwu Park, Yeonjun In, Sein Kim 외 arxiv

Although recent tool-augmented benchmarks involve complex requests, evaluation remains limited to answer matching, neglecting critical trajectory aspects like efficiency, hallucination, and adaptivity. The most straightf…

Time Series Augmented Generation for Financial Applications

2026-04-21 · Anton Kolonin, Alexey Glushchenko, Evgeny Bochkov, Abhishek Saxena arxiv

Evaluating the reasoning capabilities of Large Language Models (LLMs) for complex, quantitative financial tasks is a critical and unsolved challenge. Standard benchmarks often fail to isolate an agent's core ability to p…