paper-with-me

홈 › Papers

Improving Factuality of 3D Brain MRI Report Generation with Paired Image-domain Retrieval and Text-domain Augmentation

2024-11-23 · Junhyeok Lee, Yujin Oh, Dahyoun Lee, Hyon Keun Joh, Chul-Ho Sohn, Sung Hyun Baik, Cheol Kyu Jung, Jung Hyun Park, Kyu Sung Choi, Byung-Hoon Kim, Jong Chul Ye

Acute ischemic stroke (AIS) requires time-critical management, with hours of delayed intervention leading to an irreversible disability of the patient. Since diffusion weighted imaging (DWI) using the magnetic resonance image (MRI) plays a crucial role in the detection of AIS, automated prediction of AIS from DWI has been a research topic of clinical importance. While text radiology reports contain the most relevant clinical information from the image findings, the difficulty of mapping across different modalities has limited the factuality of conventional direct DWI-to-report generation methods. Here, we propose paired image-domain retrieval and text-domain augmentation (PIRTA), a cross-modal retrieval-augmented generation (RAG) framework for providing clinician-interpretative AIS radiology reports with improved factuality. PIRTA mitigates the need for learning cross-modal mapping, which poses difficulty in image-to-text generation, by casting the cross-modal mapping problem as an in-domain retrieval of similar DWI images that have paired ground-truth text radiology reports. By exploiting the retrieved radiology reports to augment the report generation process of the query image, we show by experiments with extensive in-house and public datasets that PIRTA can accurately retrieve relevant reports from 3D DWI images. This approach enables the generation of radiology reports with significantly higher accuracy compared to direct image-to-text generation using state-of-the-art multimodal language models.

📄 PDF Abstract BibTeX arXiv:2411.15490

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Modal RetrievalImage to textRAGRetrievalRetrieval-augmented GenerationText Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

BTReport: A Framework for Brain Tumor Radiology Report Generation with Clinically Relevant Features

2026-02-17 · Juampablo E. Heras Rivera, Dickson T. Chen, Tianyi Ren, Daniel K. Low 외 arxiv

Recent advances in radiology report generation (RRG) have been driven by large paired image-text datasets; however, progress in neuro-oncology has been limited due to a lack of open paired image-report datasets. Here, we…

Multi-LLM Collaborative MRI Report Generation for Visual Instruction Tuning in Brain Oncology

2026-07-16 · Sinyoung Ra, Jonghun Kim, Hyunjin Park arxiv

Recent advances in large language models (LLMs) and their extension to vision-language models (VLMs) have made it easier to combine text and images for tasks such as report generation. Existing VLMs in medicine typically…

Visual Question Answering

MedCycle: Unpaired Medical Report Generation via Cycle-Consistency

2024-03-20 · Elad Hirsch, Gefen Dawidowicz, Ayellet Tal

Generating medical reports for X-ray images presents a significant challenge, particularly in unpaired scenarios where access to paired image-report data for training is unavailable. Previous works have typically learned…

Medical Report Generation

MedRAT: Unpaired Medical Report Generation via Auxiliary Tasks

2024-07-04 · Elad Hirsch, Gefen Dawidowicz, Ayellet Tal

Medical report generation from X-ray images is a challenging task, particularly in an unpaired setting where paired image-report data is unavailable for training. To address this challenge, we propose a novel model that …

Contrastive LearningMedical Report Generation

Visual Instruction-Finetuned Language Model for Versatile Brain MR Image Tasks

2026-04-03 · Jonghun Kim, Sinyoung Ra, Hyunjin Park arxiv

LLMs have demonstrated remarkable capabilities in linguistic reasoning and are increasingly adept at vision-language tasks. The integration of image tokens into transformers has enabled direct visual input and output, ad…

Visual Question AnsweringText-to-Image GenerationImage SegmentationVisual Reasoning