paper-with-me

Papers

PathFound: An Agentic Multimodal Model Activating Evidence-seeking Pathological Diagnosis

2025-12-29 · Shengyi Hua, Jianfeng Wu, Tianle Shen, Kangzhe Hu, Zhongzhen Huang, Shujuan Ni, Zhihong Zhang, Yuan Li, Zhe Wang, Xiaofan Zhang arxiv

Recent pathological foundation models have substantially advanced visual representation learning and multimodal interaction. However, most models still rely on a static inference paradigm in which whole-slide images are processed once to produce predictions, without reassessment or targeted evidence acquisition under ambiguous diagnoses. This contrasts with clinical diagnostic workflows that refine hypotheses through repeated slide observations and further examination requests. We propose PathFound, an agentic multimodal model designed to support evidence-seeking inference in pathological diagnosis. PathFound integrates the power of pathological visual foundation models, vision-language models, and reasoning models trained with reinforcement learning to perform proactive information acquisition and diagnosis refinement by progressing through the initial diagnosis, evidence-seeking, and final decision stages. Across several large multimodal models, adopting this strategy consistently improves diagnostic accuracy, indicating the effectiveness of evidence-seeking workflows in computational pathology. Among these models, PathFound achieves state-of-the-art diagnostic performance across diverse clinical scenarios and demonstrates strong potential to discover subtle details, such as nuclear features and local invasions.

📄 PDF Abstract BibTeX arXiv:2512.23545

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningReinforcement Learning

Similar Papers 제목 키워드 기반

ClinSeekAgent: Automating Multimodal Evidence Seeking for Agentic Clinical Reasoning

2026-05-19 · Juncheng Wu, Letian Zhang, Yuhan Wang, Haoqin Tu 외 arxiv

Large language models (LLMs) and agentic systems have shown promise for clinical decision support, but existing works largely assume that evidence has already been curated and handed to the model. Real-world clinical wor…

InterLV-Search: Benchmarking Interleaved Multimodal Agentic Search

2026-05-08 · Bohan Hou, Jiuning Gu, Jiayan Guo, Ronghao Dang 외 arxiv

Existing benchmarks for multimodal agentic search evaluate multimodal search and visual browsing, but visual evidence is either confined to the input or treated as an answer endpoint rather than part of an interleaved se…

Struct-Searcher: Agentic Structural Thinking Advances Multimodal Deep Information Seeking

2026-06-05 · Fan Zhang, Vireo Zhang, Shengju Qian, Haoxuan Li 외 arxiv

Deep research agents have attracted increasing attention for their ability to collect large-scale online information to acquire target knowledge, with recent efforts shifting from purely text-based information seeking to…

Sparse Evidence Can Suffice: Agentic Evidence Seeking for Multimodal Video Misinformation Detection

2026-07-20 · Haochen Zhao, Yongxiu Xu, Xinkui Lin, Dong Xie 외 arxiv

Multimodal video misinformation detection is commonly formulated as a holistic video-understanding task, where the entire video and its associated content are processed and judged in a single pass. However, real-world mi…

Reinforcement LearningMultimodal Reasoning

Agentic Search in the Wild: Intents and Trajectory Dynamics from 14M+ Real Search Requests

2026-01-24 · Jingjie Ning, João Coelho, Yibo Kong, Yunfan Long 외 arxiv

LLM-powered search agents are increasingly being used for multi-step information seeking tasks, yet the IR community lacks empirical understanding of how agentic search sessions unfold and how retrieved evidence is refle…