paper-with-me

홈 › Papers

Layout-Guided Controllable Pathology Image Generation with In-Context Diffusion Transformers

2026-03-11 · Yuntao Shou, Xiangyong Cao, Qian Zhao, Deyu Meng arxiv

Controllable pathology image synthesis requires reliable regulation of spatial layout, tissue morphology, and semantic detail. However, existing text-guided diffusion models offer only coarse global control and lack the ability to enforce fine-grained structural constraints. Progress is further limited by the absence of large datasets that pair patch-level spatial layouts with detailed diagnostic descriptions, since generating such annotations for gigapixel whole-slide images is prohibitively time-consuming for human experts. To overcome these challenges, we first develop a scalable multi-agent LVLM annotation framework that integrates image description, diagnostic step extraction, and automatic quality judgment into a coordinated pipeline, and we evaluate the reliability of the system through a human verification process. This framework enables efficient construction of fine-grained and clinically aligned supervision at scale. Building on the curated data, we propose In-Context Diffusion Transformer (IC-DiT), a layout-aware generative model that incorporates spatial layouts, textual descriptions, and visual embeddings into a unified diffusion transformer. Through hierarchical multimodal attention, IC-DiT maintains global semantic coherence while accurately preserving structural and morphological details. Extensive experiments on five histopathology datasets show that IC-DiT achieves higher fidelity, stronger spatial controllability, and better diagnostic consistency than existing methods. In addition, the generated images serve as effective data augmentation resources for downstream tasks such as cancer classification and survival analysis.

📄 PDF Abstract BibTeX arXiv:2603.13386

Code (0)

등록된 구현이 없습니다.

Tasks

Cancer ClassificationData AugmentationImage Generation

Similar Papers 제목 키워드 기반

Spatial Diffusion for Cell Layout Generation

2024-09-04 · Chen Li, Xiaoling Hu, Shahira Abousamra, Meilong Xu 외

Generative models, such as GANs and diffusion models, have been used to augment training sets and boost performances in different tasks. We focus on generative models for cell detection instead, i.e., locating and classi…

Cell DetectionLayout Generation

ConsistCompose: Unified Multimodal Layout Control for Image Composition

2025-11-23 · Xuanke Shi, Boxuan Li, Xiaoyang Han, Zhongang Cai 외 arxiv

Unified multimodal models that couple visual understanding with image generation have advanced rapidly, yet most systems still focus on visual grounding-aligning language with image regions-while their generative counter…

Visual GroundingImage Generation

Diagnostic Benchmark and Iterative Inpainting for Layout-Guided Image Generation

2023-04-13 · Jaemin Cho, Linjie Li, Zhengyuan Yang, Zhe Gan 외

Spatial control is a core capability in controllable image generation. Advancements in layout-guided image generation have shown promising results on in-distribution (ID) datasets with similar spatial configurations. How…

DiagnosticImage GenerationLayout-to-Image Generation

QCAgent: An agentic framework for quality-controllable pathology report generation from whole slide image

2026-03-02 · Rundong Wang, Wei Ba, Ying Zhou, Yingtai Li 외 arxiv

Recent methods for pathology report generation from whole-slide image (WSI) are capable of producing slide-level diagnostic descriptions but fail to ground fine-grained statements in localized visual evidence. Furthermor…

Semantic Retrieval

LocRef-Diffusion:Tuning-Free Layout and Appearance-Guided Generation

2024-11-22 · Fan Deng, Yaguang Wu, Xinyang Yu, Xiangjun Huang 외

Recently, text-to-image models based on diffusion have achieved remarkable success in generating high-quality images. However, the challenge of personalized, controllable generation of instances within these images remai…