Learning Visual-Semantic Embeddings for Reporting Abnormal Findings on Chest X-rays
Automatic medical image report generation has drawn growing attention due to its potential to alleviate radiologists' workload. Existing work on report generation often trains encoder-decoder networks to generate complete reports. However, such models are affected by data bias (e.g.~label imbalance) and face common issues inherent in text generation models (e.g.~repetition). In this work, we focus on reporting abnormal findings on radiology images; instead of training on complete radiology reports, we propose a method to identify abnormal findings from the reports in addition to grouping them with unsupervised clustering and minimal rules. We formulate the task as cross-modal retrieval and propose Conditional Visual-Semantic Embeddings to align images and fine-grained abnormal findings in a joint embedding space. We demonstrate that our method is able to retrieve abnormal findings and outperforms existing generation models on both clinical correctness and text generation metrics.
Code (0)
등록된 구현이 없습니다.
Tasks
ClusteringCross-Modal RetrievalDecoderRetrievalText GenerationSimilar Papers 제목 키워드 기반
Impacts of National Cultures on Managerial Decisions of Engaging in Core Earnings Management
This study investigates the impact of Hofstede's cultural dimensions on abnormal core earnings management in multiple national cultural contexts. We employ an Ordinary Least Squares (OLS) regression model with abnormal c…
ManagementAssessing the Performance of Automated Prediction and Ranking of Patient Age from Chest X-rays Against Clinicians
Understanding the internal physiological changes accompanying the aging process is an important aspect of medical image interpretation, with the expected changes acting as a baseline when reporting abnormal findings. Dee…
Age EstimationGenerative Adversarial NetworkPredictionCXRMate-2: Structured Multimodal Temporal Embeddings and Tractable Reinforcement Learning for Clinically Acceptable Chest X-ray Radiology Report Generation
Chest X-ray (CXR) radiology report generation (RRG) models have shown rapid progress on automated metrics, yet their clinical utility remains uncertain due to limited qualitative evaluation by radiologists. We present CX…
Reinforcement LearningBoosting Vision Semantic Density with Anatomy Normality Modeling for Medical Vision-language Pre-training
Vision-language pre-training (VLP) has great potential for developing multifunctional and general medical diagnostic capabilities. However, aligning medical images with a low signal-to-noise ratio (SNR) to reports with a…
Contrastive LearningTransfer LearningMedVista3D: Vision-Language Modeling for Reducing Diagnostic Errors in 3D CT Disease Detection, Understanding and Reporting
Radiologic diagnostic errors-under-reading errors, inattentional blindness, and communication failures-remain prevalent in clinical practice. These issues often stem from missed localized abnormalities, limited global co…
Visual Question AnsweringRepresentation LearningSpatial Reasoning