paper-with-me

Papers

Contextualized Keyword Representations for Multi-modal Retinal Image Captioning

2021-04-26 · Jia-Hong Huang, Ting-Wei Wu, Marcel Worring

Medical image captioning automatically generates a medical description to describe the content of a given medical image. A traditional medical image captioning model creates a medical description only based on a single medical image input. Hence, an abstract medical description or concept is hard to be generated based on the traditional approach. Such a method limits the effectiveness of medical image captioning. Multi-modal medical image captioning is one of the approaches utilized to address this problem. In multi-modal medical image captioning, textual input, e.g., expert-defined keywords, is considered as one of the main drivers of medical description generation. Thus, encoding the textual input and the medical image effectively are both important for the task of multi-modal medical image captioning. In this work, a new end-to-end deep multi-modal medical image captioning model is proposed. Contextualized keyword representations, textual feature reinforcement, and masked self-attention are used to develop the proposed approach. Based on the evaluation of the existing multi-modal medical image captioning dataset, experimental results show that the proposed model is effective with the increase of +53.2% in BLEU-avg and +18.6% in CIDEr, compared with the state-of-the-art method.

📄 PDF Abstract BibTeX arXiv:2104.12471

Code (0)

등록된 구현이 없습니다.

Tasks

AvgImage Captioning

Similar Papers 제목 키워드 기반

M3T: Multi-Modal Medical Transformer to bridge Clinical Context with Visual Insights for Retinal Image Medical Description Generation

2024-06-19 · Nagur Shareef Shaik, Teja Krishna Cherukuri, Dong Hye Ye

Automated retinal image medical description generation is crucial for streamlining medical diagnosis and treatment planning. Existing challenges include the reliance on learned retinal image representations, difficulties…

DiagnosticMedical Diagnosis

DREAM: Dynamic Retinal Enhancement with Adaptive Multi-modal Fusion for Expert Precision Medical Report Generation

2026-04-19 · Nagur Shareef Shaik, Teja Krishna Cherukuri, Dong Hye Ye arxiv

Automating medical reports for retinal images requires a sophisticated blend of visual pattern recognition and deep clinical knowledge. Current Large Vision-Language Models (LVLMs) often struggle in specialized medical f…

Medical Report GenerationClinical Knowledge

Longer Version for "Deep Context-Encoding Network for Retinal Image Captioning"

2021-05-30 · Jia-Hong Huang, Ting-Wei Wu, Chao-Han Huck Yang, Marcel Worring

Automatically generating medical reports for retinal images is one of the promising ways to help ophthalmologists reduce their workload and improve work efficiency. In this work, we propose a new context-driven encoding …

AvgDecoderImage CaptioningMedical Report Generation

Automated Retinal Image Analysis and Medical Report Generation through Deep Learning

2024-08-14 · Jia-Hong Huang

The increasing prevalence of retinal diseases poses a significant challenge to the healthcare system, as the demand for ophthalmologists surpasses the available workforce. This imbalance creates a bottleneck in diagnosis…

DiagnosticMedical Report Generation

Contextualized Weak Supervision for Text Classification

2020-07-01 · ACL 2020 6 · Dheeraj Mekala, Jingbo Shang

Weakly supervised text classification based on a few user-provided seed words has recently attracted much attention from researchers. Existing methods mainly generate pseudo-labels in a context-free manner (e.g., string …

ClassificationGeneral Classificationtext-classificationText Classification