paper-with-me

Papers

RetinaLogos: Fine-Grained Synthesis of High-Resolution Retinal Images Through Captions

2025-05-19 · Junzhi Ning, Cheng Tang, Kaijin Zhou, Diping Song, Lihao Liu, Ming Hu, Wei Li, Yanzhou Su, Tianbing Li, Jiyao Liu, Yejin, Sheng Zhang, Yuanfeng Ji, Junjun He

The scarcity of high-quality, labelled retinal imaging data, which presents a significant challenge in the development of machine learning models for ophthalmology, hinders progress in the field. To synthesise Colour Fundus Photographs (CFPs), existing methods primarily relying on predefined disease labels face significant limitations. However, current methods remain limited, thus failing to generate images for broader categories with diverse and fine-grained anatomical structures. To overcome these challenges, we first introduce an innovative pipeline that creates a large-scale, synthetic Caption-CFP dataset comprising 1.4 million entries, called RetinaLogos-1400k. Specifically, RetinaLogos-1400k uses large language models (LLMs) to describe retinal conditions and key structures, such as optic disc configuration, vascular distribution, nerve fibre layers, and pathological features. Furthermore, based on this dataset, we employ a novel three-step training framework, called RetinaLogos, which enables fine-grained semantic control over retinal images and accurately captures different stages of disease progression, subtle anatomical variations, and specific lesion types. Extensive experiments demonstrate state-of-the-art performance across multiple datasets, with 62.07% of text-driven synthetic images indistinguishable from real ones by ophthalmologists. Moreover, the synthetic data improves accuracy by 10%-25% in diabetic retinopathy grading and glaucoma detection, thereby providing a scalable solution to augment ophthalmic datasets.

📄 PDF Abstract BibTeX arXiv:2505.12887

Code (0)

등록된 구현이 없습니다.

Tasks

Diabetic Retinopathy Grading

Similar Papers 제목 키워드 기반

Fine-grained Semantic Constraint in Image Synthesis

2021-01-12 · Pengyang Li, Donghui Wang

In this paper, we propose a multi-stage and high-resolution model for image synthesis that uses fine-grained attributes and masks as input. With a fine-grained attribute, the proposed model can detailedly constrain the f…

AttributeDiversityGenerative Adversarial NetworkImage Generation

Fine-grained Cross-modal Fusion based Refinement for Text-to-Image Synthesis

2023-02-17 · Haoran Sun, Yang Wang, Haipeng Liu, Biao Qian

Text-to-image synthesis refers to generating visual-realistic and semantically consistent images from given textual descriptions. Previous approaches generate an initial low-resolution image and then refine it to be high…

Image Generation

Hierarchical Multi-Grained Generative Model for Expressive Speech Synthesis

2020-09-17 · Yukiya Hono, Kazuna Tsuboi, Kei Sawada, Kei Hashimoto 외

This paper proposes a hierarchical generative model with a multi-grained latent variable to synthesize expressive speech. In recent years, fine-grained latent variables are introduced into the text-to-speech synthesis th…

Expressive Speech SynthesisSpeech Synthesistext-to-speechText to Speech+1

Latent Wavelet Diffusion: Enabling 4K Image Synthesis for Free

2025-05-31 · Luigi Sigillo, Shengfeng He, Danilo Comminiello

High-resolution image synthesis remains a core challenge in generative modeling, particularly in balancing computational efficiency with the preservation of fine-grained visual detail. We present Latent Wavelet Diffusion…

2k4kComputational EfficiencyDenoising+1

ViBe: Ultra-High-Resolution Video Synthesis Born from Pure Images

2026-03-24 · Yunfeng Wu, Hongying Cheng, Zihao He, Songhua Liu arxiv

Transformer-based video diffusion models rely on 3D attention over spatial and temporal tokens, which incurs quadratic time and memory complexity and makes end-to-end training for ultra-high-resolution videos prohibitive…

Video Generation