paper-with-me

Papers

StableSketcher: Enhancing Diffusion Model for Pixel-based Sketch Generation via Visual Question Answering Feedback

2025-10-23 · Jiho Park, Sieun Choi, Jaeyoon Seo, Jihie Kim arxiv

Although recent advancements in diffusion models have significantly enriched the quality of generated images, challenges remain in synthesizing pixel-based human-drawn sketches, a representative example of abstract expression. To combat these challenges, we propose StableSketcher, a novel framework that empowers diffusion models to generate hand-drawn sketches with high prompt fidelity. Within this framework, we fine-tune the variational autoencoder to optimize latent decoding, enabling it to better capture the characteristics of sketches. In parallel, we integrate a new reward function for reinforcement learning based on visual question answering, which improves text-image alignment and semantic consistency. Extensive experiments demonstrate that StableSketcher generates sketches with improved stylistic fidelity, achieving better alignment with prompts compared to the Stable Diffusion baseline. Additionally, we introduce SketchDUO, to the best of our knowledge, the first dataset comprising instance-level sketches paired with captions and question-answer pairs, thereby addressing the limitations of existing datasets that rely on image-label pairs. Our code and dataset will be made publicly available upon acceptance. Project page: https://zihos.github.io/StableSketcher

📄 PDF Abstract BibTeX arXiv:2510.20093

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringReinforcement Learning

Similar Papers 제목 키워드 기반

CoProSketch: Controllable and Progressive Sketch Generation with Diffusion Model

2025-04-11 · Ruohao Zhan, Yijin Li, Yisheng He, Shuo Chen 외

Sketches serve as fundamental blueprints in artistic creation because sketch editing is easier and more intuitive than pixel-level RGB image editing for painting artists, yet sketch generation remains unexplored despite …

Image Generation

Sketch-Guided Motion Diffusion for Stylized Cinemagraph Synthesis

2024-12-01 · Hao Jin, Hengyuan Chang, Xiaoxuan Xie, Zhengyang Wang 외

Designing stylized cinemagraphs is challenging due to the difficulty in customizing complex and expressive flow motions. To achieve intuitive and detailed control of the generated cinemagraphs, freehand sketches can prov…

object-detectionObject Detection

CLIP4Sketch: Enhancing Sketch to Mugshot Matching through Dataset Augmentation using Diffusion Models

2024-08-02 · Kushal Kumar Jain, Steve Grosz, Anoop M. Namboodiri, Anil K. Jain

Forensic sketch-to-mugshot matching is a challenging task in face recognition, primarily hindered by the scarcity of annotated forensic sketches and the modality gap between sketches and photographs. To address this, we …

DenoisingFace Recognition

Inversion-by-Inversion: Exemplar-based Sketch-to-Photo Synthesis via Stochastic Differential Equations without Training

2023-08-15 · XiMing Xing, Chuang Wang, Haitao Zhou, Zhihao Hu 외

Exemplar-based sketch-to-photo synthesis allows users to generate photo-realistic images based on sketches. Recently, diffusion-based methods have achieved impressive performance on image generation tasks, enabling highl…

Image Generation

SketchDreamer: Interactive Text-Augmented Creative Sketch Ideation

2023-08-27 · Zhiyu Qu, Tao Xiang, Yi-Zhe Song

Artificial Intelligence Generated Content (AIGC) has shown remarkable progress in generating realistic images. However, in this paper, we take a step "backward" and address AIGC for the most rudimentary visual modality o…