paper-with-me

홈 › Papers

Self-Supervised Anatomical Consistency Learning for Vision-Grounded Medical Report Generation

2025-09-30 · Longzhen Yang, Zhangkai Ni, Ying Wen, Yihang Liu, Lianghua He, Heng Tao Shen arxiv

Vision-grounded medical report generation aims to produce clinically accurate descriptions of medical images, anchored in explicit visual evidence to improve interpretability and facilitate integration into clinical workflows. However, existing methods often rely on separately trained detection modules that require extensive expert annotations, introducing high labeling costs and limiting generalizability due to pathology distribution bias across datasets. To address these challenges, we propose Self-Supervised Anatomical Consistency Learning (SS-ACL) -- a novel and annotation-free framework that aligns generated reports with corresponding anatomical regions using simple textual prompts. SS-ACL constructs a hierarchical anatomical graph inspired by the invariant top-down inclusion structure of human anatomy, organizing entities by spatial location. It recursively reconstructs fine-grained anatomical regions to enforce intra-sample spatial alignment, inherently guiding attention maps toward visually relevant areas prompted by text. To further enhance inter-sample semantic alignment for abnormality recognition, SS-ACL introduces a region-level contrastive learning based on anatomical consistency. These aligned embeddings serve as priors for report generation, enabling attention maps to provide interpretable visual evidence. Extensive experiments demonstrate that SS-ACL, without relying on expert annotations, (i) generates accurate and visually grounded reports -- outperforming state-of-the-art methods by 10\% in lexical accuracy and 25\% in clinical efficacy, and (ii) achieves competitive performance on various downstream visual tasks, surpassing current leading visual foundation models by 8\% in zero-shot visual grounding.

📄 PDF Abstract BibTeX arXiv:2509.25963

Code (0)

등록된 구현이 없습니다.

Tasks

Medical Report GenerationContrastive LearningVisual Grounding

Similar Papers 제목 키워드 기반

AFiRe: Anatomy-Driven Self-Supervised Learning for Fine-Grained Representation in Radiographic Images

2025-04-15 · Yihang Liu, Lianghua He, Ying Wen, Longzhen Yang 외

Current self-supervised methods, such as contrastive learning, predominantly focus on global discrimination, neglecting the critical fine-grained anatomical details required for accurate radiographic analysis. To address…

AnatomyAnomaly DetectionContrastive LearningMulti-Label Classification+2

Learning Anatomically Consistent Embedding for Chest Radiography

2023-12-01 · Ziyu Zhou, Haozhe Luo, Jiaxuan Pang, Xiaowei Ding 외

Self-supervised learning (SSL) approaches have recently shown substantial success in learning visual representations from unannotated images. Compared with photographic images, medical images acquired with the same imagi…

AnatomyMedical Image AnalysisSelf-Supervised Learning

Transferable Visual Words: Exploiting the Semantics of Anatomical Patterns for Self-supervised Learning

2021-02-21 · Fatemeh Haghighi, Mohammad Reza Hosseinzadeh Taher, Zongwei Zhou, Michael B. Gotway 외

This paper introduces a new concept called "transferable visual words" (TransVW), aiming to achieve annotation efficiency for deep learning in medical image analysis. Medical imaging--focusing on particular parts of the …

AnatomyMedical Image AnalysisMedical Image SegmentationSelf-Supervised Learning+1

PULSE: A Unified Multi-Task Architecture for Cardiac Segmentation, Diagnosis, and Few-Shot Cross-Modality Clinical Adaptation

2025-12-03 · Hania Ghouse, Maryam Alsharqi, Farhad R. Nezami, Muzammil Behzad arxiv

Cardiac image analysis remains fragmented across tasks: anatomical segmentation, disease classification, and grounded clinical report generation are typically handled by separate networks trained under different data reg…

VLRC: Vision-Language Reprojection Consistency as a scalable signal for better feed-forward 3D pretraining

2026-07-02 · Marwane Hariat, David Filliat, Antoine Manzanera arxiv

Feed-forward 3D models are commonly trained using either expensive geometric supervision or self-supervised photometric objectives, both of which provide incomplete learning signals. We introduce Vision-Language Reprojec…

3D Semantic SegmentationScene Understanding3D Reconstruction