paper-with-me

홈 › Papers

U-VLM: Hierarchical Vision Language Modeling for Report Generation

2026-02-28 · Pengcheng Shi, Minghui Zhang, Kehan Song, Jiaqi Liu, Yun Gu, Xinglin Zhang arxiv

Automated radiology report generation is key for reducing radiologist workload and improving diagnostic consistency, yet generating accurate reports for 3D medical imaging remains challenging. Existing vision-language models face two limitations: they do not leverage segmentation-pretrained encoders, and they inject visual features only at the input layer of language models, losing multi-scale information. We propose U-VLM, which enables hierarchical vision-language modeling in both training and architecture: (1) progressive training from segmentation to classification to report generation, and (2) multi-layer visual injection that routes U-Net encoder features to corresponding language model layers. Each training stage can leverage different datasets without unified annotations. U-VLM achieves state-of-the-art performance on CT-RATE (F1: 0.414 vs 0.258, BLEU-mean: 0.349 vs 0.305) and AbdomenAtlas 3.0 (F1: 0.624 vs 0.518 for segmentation-based detection) using only a 0.1B decoder trained from scratch, demonstrating that well-designed vision encoder pretraining outweighs the benefits of 7B+ pre-trained language models. Ablation studies show that progressive pretraining significantly improves F1, while multi-layer injection improves BLEU-mean. Code is available at https://github.com/yinghemedical/U-VLM.

📄 PDF Abstract BibTeX arXiv:2603.00479

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AHIVE: Anatomy-aware Hierarchical Vision Encoding for Interactive Radiology Report Retrieval

2024-01-01 · CVPR 2024 1 · Sixing Yan, William K. Cheung, Ivor W. Tsang, Keith Chiu 외

Automatic radiology report generation using deep learning models has been recently explored and found promising. Neural decoders are commonly used for the report generation where irrelevant and unfaithful contents ar…

AnatomyDiagnosticRetrieval

RadHiera: Semantic Hierarchical Reinforcement Learning for Medical Report Generation

2025-11-13 · Bodong Du, Honglong Yang, Xiaomeng Li arxiv

Vision-language models have shown promising results in radiology report generation. However, most existing methods generate reports as flat text and do not explicitly model the semantic dependency between the Findings an…

Hierarchical Reinforcement LearningMedical Report Generation

R2GenKG: Hierarchical Multi-modal Knowledge Graph for LLM-based Radiology Report Generation

2025-08-05 · Futian Wang, Yuhan Qiao, Xiao Wang, Fuling Wang 외 arxiv

X-ray medical report generation is one of the important applications of artificial intelligence in healthcare. With the support of large foundation models, the quality of medical report generation has significantly impro…

Medical Report Generation

Writing by Memorizing: Hierarchical Retrieval-based Medical Report Generation

2021-05-25 · ACL 2021 5 · Xingyi Yang, Muchao Ye, Quanzeng You, Fenglong Ma

Medical report generation is one of the most challenging tasks in medical image analysis. Although existing approaches have achieved promising results, they either require a predefined template database in order to retri…

DecoderMedical Image AnalysisMedical Report GenerationRetrieval+1

ResTok: Learning Hierarchical Residuals in 1D Visual Tokenizers for Autoregressive Image Generation

2026-01-07 · Xu Zhang, Cheng Da, Huan Yang, Kun Gai 외 arxiv

Existing 1D visual tokenizers for autoregressive (AR) generation largely follow the design principles of language modeling, as they are built directly upon transformers whose priors originate in language, yielding single…

Image Generation