paper-with-me

홈 › Papers

SketchFlex: Facilitating Spatial-Semantic Coherence in Text-to-Image Generation with Region-Based Sketches

2025-02-11 · Haichuan Lin, Yilin Ye, Jiazhi Xia, Wei Zeng

Text-to-image models can generate visually appealing images from text descriptions. Efforts have been devoted to improving model controls with prompt tuning and spatial conditioning. However, our formative study highlights the challenges for non-expert users in crafting appropriate prompts and specifying fine-grained spatial conditions (e.g., depth or canny references) to generate semantically cohesive images, especially when multiple objects are involved. In response, we introduce SketchFlex, an interactive system designed to improve the flexibility of spatially conditioned image generation using rough region sketches. The system automatically infers user prompts with rational descriptions within a semantic space enriched by crowd-sourced object attributes and relationships. Additionally, SketchFlex refines users' rough sketches into canny-based shape anchors, ensuring the generation quality and alignment of user intentions. Experimental results demonstrate that SketchFlex achieves more cohesive image generations than end-to-end models, meanwhile significantly reducing cognitive load and better matching user intentions compared to region-based generation baseline.

📄 PDF Abstract BibTeX arXiv:2502.07556

Code (1)

SellLin/SketchFlex 공식 구현 pytorch

Tasks

Image GenerationText to Image GenerationText-to-Image Generation

Similar Papers 제목 키워드 기반

FuseNet: Self-Supervised Dual-Path Network for Medical Image Segmentation

2023-11-22 · Amirhossein Kazerouni, Sanaz Karimijafarbigloo, Reza Azad, Yury Velichko 외

Semantic segmentation, a crucial task in computer vision, often relies on labor-intensive and costly annotated datasets for training. In response to this challenge, we introduce FuseNet, a dual-stream framework for self-…

Image SegmentationMedical Image SegmentationSegmentationSemantic Segmentation

ProxyCLIP: Proxy Attention Improves CLIP for Open-Vocabulary Segmentation

2024-08-09 · Mengcheng Lan, Chaofeng Chen, Yiping Ke, Xinjiang Wang 외

Open-vocabulary semantic segmentation requires models to effectively integrate visual representations with open-vocabulary semantic labels. While Contrastive Language-Image Pre-training (CLIP) models shine in recognizing…

Open Vocabulary Semantic SegmentationOpen-Vocabulary Semantic SegmentationSegmentationSemantic Segmentation+1

GSTran: Joint Geometric and Semantic Coherence for Point Cloud Segmentation

2024-08-21 · Abiao Li, Chenlei Lv, Guofeng Mei, Yifan Zuo 외

Learning meaningful local and global information remains a challenge in point cloud segmentation tasks. When utilizing local information, prior studies indiscriminately aggregates neighbor information from different clas…

Point Cloud SegmentationSemantic SimilaritySemantic Textual Similarity

Enhancing Low-Cost Video Editing with Lightweight Adaptors and Temporal-Aware Inversion

2025-01-08 · Yangfan He, Sida Li, Kun Li, Xinyuan Song 외

Recent advancements in text-to-image (T2I) generation using diffusion models have enabled cost-effective video-editing applications by leveraging pre-trained models, eliminating the need for resource-intensive training. …

Video Editing

Plug-in Feedback Self-adaptive Attention in CLIP for Training-free Open-Vocabulary Segmentation

2025-08-27 · Zhixiang Chi, Yanan Wu, Li Gu, Huan Liu 외 arxiv

CLIP exhibits strong visual-textual alignment but struggle with open-vocabulary segmentation due to poor localization. Prior methods enhance spatial coherence by modifying intermediate attention. But, this coherence isn'…