paper-with-me

홈 › Papers

Semantic Context-aware mOdality fUsion Transformer (SCOUT): A Context-Aware Multimodal Transformer for Concept-Grounded Pathology Report Generation

2026-05-01 · Suryakant Singh, Saarthak Kapse, Joel Saltz, Prateek Prasanna arxiv

Whole-slide images (WSIs) present a fundamental challenge for computational pathology due to their extreme resolution, multi-scale heterogeneity, and the requirement for clinically reliable interpretation. Although recent pathology foundation models have enabled fluent report generation, they often lack clinical grounding, failing to accurately represent key diagnostic concepts and relationships observed by pathologists. This limitation arises from the difficulty of integrating heterogeneous visual evidence spanning fine-grained cellular patterns, slide-level tissue architecture, and high-level diagnostic concepts, while maintaining interpretability and clinical coherence. Here we present SCOUT: Semantic Context-aware mOdality fUsion Transformer, a context-aware concept-grounded multimodal framework for pathology report generation that enables progressive conditioning of image representations by global slide information and explicit diagnostic concepts. The method integrates local histological patterns, whole-slide context, and expert-curated semantic descriptors within a unified learning paradigm, allowing visual features to be dynamically refined throughout the encoding process. By combining depth-aware contextual modulation with adaptive multimodal fusion during text generation, the framework produces clinically coherent reports while preserving complementarity across representational scales. Using CONCH1.5 features, we evaluate SCOUT against WSI-Caption, HistGen, and BiGen on TCGA-BRCA, MICCAI REG, and HistAI. SCOUT achieves the best BLEU-1 to BLEU-4 and METEOR scores on all datasets, plus the best ROUGE-L on TCGA-BRCA and MICCAI REG. On TCGA-BRCA, it reaches 0.436/0.303/0.202/0.156 BLEU-1/2/3/4 and 0.204 METEOR; on REG 2025, it achieves 0.865/0.834/0.805/0.780 and 0.568. These results support progressive contextual conditioning for grounded pathology report generation.

📄 PDF Abstract BibTeX arXiv:2605.01144

Code (0)

등록된 구현이 없습니다.

Tasks

Text Generation

Similar Papers 제목 키워드 기반

Similarity Guided Multimodal Fusion Transformer for Semantic Location Prediction in Social Media

2024-05-09 · Zhizhen Zhang, Ning Wang, Haojie Li, Zhihui Wang

Semantic location prediction aims to derive meaningful location insights from multimodal social media posts, offering a more contextual understanding of daily activities than using GPS coordinates. This task faces signif…

Language Modelling

MyGram: Modality-aware Graph Transformer with Global Distribution for Multi-modal Entity Alignment

2026-01-17 · Zhifei Li, Ziyue Qin, Xiangyu Luo, Xiaoju Hou 외 arxiv

Multi-modal entity alignment aims to identify equivalent entities between two multi-modal Knowledge graphs by integrating multi-modal data, such as images and text, to enrich the semantic representations of entities. How…

Multi-modal Entity AlignmentKnowledge Graphs

NestedFormer: Nested Modality-Aware Transformer for Brain Tumor Segmentation

2022-08-31 · Zhaohu Xing, Lequan Yu, Liang Wan, Tong Han 외

Multi-modal MR imaging is routinely used in clinical practice to diagnose and investigate brain tumors by providing rich complementary information. Previous multi-modal MRI segmentation methods usually perform modal fusi…

Brain Tumor SegmentationDecoderMRI segmentationSegmentation+1

SwinNet: Swin Transformer drives edge-aware RGB-D and RGB-T salient object detection

2022-04-12 · Zhengyi Liu, Yacheng Tan, Qian He, Yun Xiao

Convolutional neural networks (CNNs) are good at extracting contexture features within certain receptive fields, while transformers can model the global long-range dependency features. By absorbing the advantage of trans…

Decoderobject-detectionObject DetectionRGB-T Salient Object Detection+1

Local-to-Global Cross-Modal Attention-Aware Fusion for HSI-X Semantic Segmentation

2024-06-25 · Xuming Zhang, Naoto Yokoya, Xingfa Gu, Qingjiu Tian 외

Hyperspectral image (HSI) classification has recently reached its performance bottleneck. Multimodal data fusion is emerging as a promising approach to overcome this bottleneck by providing rich complementary information…

DecoderSemantic Segmentation