paper-with-me

Papers

Pathological Visual Question Answering

2020-10-06 · Xuehai He, Zhuo Cai, Wenlan Wei, Yichen Zhang, Luntian Mou, Eric Xing, Pengtao Xie

Is it possible to develop an "AI Pathologist" to pass the board-certified examination of the American Board of Pathology (ABP)? To build such a system, three challenges need to be addressed. First, we need to create a visual question answering (VQA) dataset where the AI agent is presented with a pathology image together with a question and is asked to give the correct answer. Due to privacy concerns, pathology images are usually not publicly available. Besides, only well-trained pathologists can understand pathology images, but they barely have time to help create datasets for AI research. The second challenge is: since it is difficult to hire highly experienced pathologists to create pathology visual questions and answers, the resulting pathology VQA dataset may contain errors. Training pathology VQA models using these noisy or even erroneous data will lead to problematic models that cannot generalize well on unseen images. The third challenge is: the medical concepts and knowledge covered in pathology question-answer (QA) pairs are very diverse while the number of QA pairs available for modeling training is limited. How to learn effective representations of diverse medical concepts based on limited data is technically demanding. In this paper, we aim to address these three challenges. To our best knowledge, our work represents the first one addressing the pathology VQA problem. To deal with the issue that a publicly available pathology VQA dataset is lacking, we create PathVQA dataset. To address the second challenge, we propose a learning-by-ignoring approach. To address the third challenge, we propose to use cross-modal self-supervised learning. We perform experiments on our created PathVQA dataset and the results demonstrate the effectiveness of our proposed learning-by-ignoring method and cross-modal self-supervised learning methods.

📄 PDF Abstract BibTeX arXiv:2010.12435

Code (0)

등록된 구현이 없습니다.

Tasks

AI AgentQuestion AnsweringSelf-Supervised LearningVisual Question AnsweringVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

American 설명 없음

Similar Papers 제목 키워드 기반

TM-PATHVQA:90000+ Textless Multilingual Questions for Medical Visual Question Answering

2024-07-16 · Tonmoy Rajkhowa, Amartya Roy Chowdhury, Sankalp Nagaonkar, Achyut Mani Tripathi

In healthcare and medical diagnostics, Visual Question Answering (VQA) mayemergeasapivotal tool in scenarios where analysis of intricate medical images becomes critical for accurate diagnoses. Current text-based VQA syst…

Medical Visual Question AnsweringQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Towards Visual Question Answering on Pathology Images

2021-08-01 · ACL 2021 5 · Xuehai He, Zhuo Cai, Wenlan Wei, Yichen Zhang 외

Pathology imaging is broadly used for identifying the causes and effects of diseases or injuries. Given a pathology image, being able to answer questions about the clinical findings contained in the image is very importa…

Decision MakingQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

WSI-VQA: Interpreting Whole Slide Images by Generative Visual Question Answering

2024-07-08 · Pingyi Chen, Chenglu Zhu, Sunyi Zheng, Honglin Li 외

Whole slide imaging is routinely adopted for carcinoma diagnosis and prognosis. Abundant experience is required for pathologists to achieve accurate and reliable diagnostic results of whole slide images (WSI). The huge s…

DiagnosticGenerative Visual Question AnsweringPrognosisQuestion Answering+5

Enhancing Pathological VLMs with Cross-scale Reasoning

2026-06-16 · Chi Phan, Tianyi Zhang, Qiaochu Xue, Yufeng Wu 외 arxiv

Pathological images are inherently multi-scale, requiring pathologists to integrate evidence from global tissue architecture at low magnification to cellular morphology at higher magnification for accurate diagnosis. Whi…

Visual Question AnsweringReinforcement Learning

Design as Desired: Utilizing Visual Question Answering for Multimodal Pre-training

2024-03-30 · Tongkun Su, Jun Li, Xi Zhang, Haibo Jin 외

Multimodal pre-training demonstrates its potential in the medical domain, which learns medical visual representations from paired medical reports. However, many pre-training tasks require extra annotations from clinician…

Contrastive LearningQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)