paper-with-me

홈 › Papers

Neural Scene Designer: Self-Styled Semantic Image Manipulation

2025-09-01 · Jianman Lin, Tianshui Chen, Chunmei Qing, Zhijing Yang, Shuangping Huang, Yuheng Ren, Liang Lin arxiv

Maintaining stylistic consistency is crucial for the cohesion and aesthetic appeal of images, a fundamental requirement in effective image editing and inpainting. However, existing methods primarily focus on the semantic control of generated content, often neglecting the critical task of preserving this consistency. In this work, we introduce the Neural Scene Designer (NSD), a novel framework that enables photo-realistic manipulation of user-specified scene regions while ensuring both semantic alignment with user intent and stylistic consistency with the surrounding environment. NSD leverages an advanced diffusion model, incorporating two parallel cross-attention mechanisms that separately process text and style information to achieve the dual objectives of semantic control and style consistency. To capture fine-grained style representations, we propose the Progressive Self-style Representational Learning (PSRL) module. This module is predicated on the intuitive premise that different regions within a single image share a consistent style, whereas regions from different images exhibit distinct styles. The PSRL module employs a style contrastive loss that encourages high similarity between representations from the same image while enforcing dissimilarity between those from different images. Furthermore, to address the lack of standardized evaluation protocols for this task, we establish a comprehensive benchmark. This benchmark includes competing algorithms, dedicated style-related metrics, and diverse datasets and settings to facilitate fair comparisons. Extensive experiments conducted on our benchmark demonstrate the effectiveness of the proposed framework.

📄 PDF Abstract BibTeX arXiv:2509.01405

Code (0)

등록된 구현이 없습니다.

Tasks

Image ManipulationImage Editing

Similar Papers 제목 키워드 기반

SemStyle: Learning to Generate Stylised Image Captions using Unaligned Text

2018-05-18 · CVPR 2018 6 · Alexander Mathews, Lexing Xie, Xuming He

Linguistic style is an essential part of written communication, with the power to affect both clarity and attractiveness. With recent advances in vision and language, we can start to tackle the problem of generating imag…

DescriptiveImage CaptioningLanguage ModelingLanguage Modelling

StyleDiff: Attribute Comparison Between Unlabeled Datasets in Latent Disentangled Space

2023-03-09 · Keisuke Kawano, Takuro Kutsuna, Ryoko Tokuhisa, Akihiro Nakamura 외

One major challenge in machine learning applications is coping with mismatches between the datasets used in the development and those obtained in real-world applications. These mismatches may lead to inaccurate predictio…

Attribute

StyleDEM: a Versatile Model for Authoring Terrains

2023-04-19 · Simon Perche, Adrien Peytavie, Bedrich Benes, Eric Galin 외

Many terrain modelling methods have been proposed for the past decades, providing efficient and often interactive authoring tools. However, they generally do not include any notion of style, which is a critical aspect fo…

Generative Adversarial NetworkmodelSuper-Resolution

FedGAI: Federated Style Learning with Cloud-Edge Collaboration for Generative AI in Fashion Design

2025-03-16 · Mingzhu Wu, Jianan Jiang, Xinglin Li, Hanhui Deng 외

Collaboration can amalgamate diverse ideas, styles, and visual elements, fostering creativity and innovation among different designers. In collaborative design, sketches play a pivotal role as a means of expressing desig…

Federated Learning

Handwriting Transformers

2021-04-08 · ICCV 2021 10 · Ankan Kumar Bhunia, Salman Khan, Hisham Cholakkal, Rao Muhammad Anwer 외

We propose a novel transformer-based styled handwritten text image generation approach, HWT, that strives to learn both style-content entanglement as well as global and local writing style patterns. The proposed HWT capt…

DecoderImage GenerationText Generation