paper-with-me

홈 › Papers

Efficient Multi-Slide Visual-Language Feature Fusion for Placental Disease Classification

2025-08-05 · Hang Guo, Qing Zhang, Zixuan Gao, Siyuan Yang, Shulin Peng, Xiang Tao, Ting Yu, Yan Wang, Qingli Li arxiv

Accurate prediction of placental diseases via whole slide images (WSIs) is critical for preventing severe maternal and fetal complications. However, WSI analysis presents significant computational challenges due to the massive data volume. Existing WSI classification methods encounter critical limitations: (1) inadequate patch selection strategies that either compromise performance or fail to sufficiently reduce computational demands, and (2) the loss of global histological context resulting from patch-level processing approaches. To address these challenges, we propose an Efficient multimodal framework for Patient-level placental disease Diagnosis, named EmmPD. Our approach introduces a two-stage patch selection module that combines parameter-free and learnable compression strategies, optimally balancing computational efficiency with critical feature preservation. Additionally, we develop a hybrid multimodal fusion module that leverages adaptive graph learning to enhance pathological feature representation and incorporates textual medical reports to enrich global contextual understanding. Extensive experiments conducted on both a self-constructed patient-level Placental dataset and two public datasets demonstrating that our method achieves state-of-the-art diagnostic performance. The code is available at https://github.com/ECNU-MultiDimLab/EmmPD.

📄 PDF Abstract BibTeX arXiv:2508.03277

Code (0)

등록된 구현이 없습니다.

Tasks

Computational EfficiencyGraph Learning

Similar Papers 제목 키워드 기반

LoC-Path: Learning to Compress for Pathology Multimodal Large Language Models

2025-12-05 · Qingqiao Hu, Weimin Lyu, Meilong Xu, Kehan Qi 외 arxiv

Whole Slide Image (WSI) MLLMs are difficult to build and deploy because gigapixel slides induce thousands of visual tokens, while only a small fraction of regions is diagnostically relevant. Existing slide-level patholog…

SliderSpace: Decomposing the Visual Capabilities of Diffusion Models

2025-02-03 · Rohit Gandikota, Zongze Wu, Richard Zhang, David Bau 외

We present SliderSpace, a framework for automatically decomposing the visual capabilities of diffusion models into controllable and human-understandable directions. Unlike existing control methods that require a user to …

Diversity

What's the Best Way to Retrieve Slides? A Comparative Study of Multimodal, Caption-Based, and Hybrid Retrieval Techniques

2025-09-18 · Petros Stylianos Giouroukis, Dimitris Dimitriadis, Dimitrios Papadopoulos, Zhenwen Shao 외 arxiv

Slide decks, serving as digital reports that bridge the gap between presentation slides and written documents, are a prevalent medium for conveying information in both academic and corporate settings. Their multimodal na…

Concept Sliders: LoRA Adaptors for Precise Control in Diffusion Models

2023-11-20 · Rohit Gandikota, Joanna Materzynska, Tingrui Zhou, Antonio Torralba 외

We present a method to create interpretable concept sliders that enable precise control over attributes in image generations from diffusion models. Our approach identifies a low-rank parameter direction corresponding to …

Image Generation

A Multi-Source Data Fusion-based Semantic Segmentation Model for Relic Landslide Detection

2023-08-02 · Yiming Zhou, Yuexing Peng, Junchuan Yu, Daqing Ge 외

As a natural disaster, landslide often brings tremendous losses to human lives, so it urgently demands reliable detection of landslide risks. When detecting relic landslides that present important information for landsli…

Contrastive LearningLandslide segmentationSemantic Segmentation