paper-with-me

Papers Image Generation

“Image Generation” 태그가 달린 논문 7,946편 · 필터 해제

Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation

2026-09-08 · Igor Pavlovic, Thiemo Wandel, Anton Obukhov, Luca Bartolomei 외 hf

Monocular depth estimation is a ubiquitous yet highly ill-posed computer vision task, with downstream applications in scene reconstruction, computational photography, and robotics, among others. Despite the field's matur…

Surface Normals EstimationMonocular Depth EstimationImage Generation

WeAgent-MMGenEdit: A Full-Stack Recipe for Multimodal Agentic Image Generation and Editing

2026-09-04 · Hui Zhang, Zongkai Liu, Liqiang Niu, Juntao Liu 외 arxiv

Image generation and editing models have advanced rapidly, yet remain unreliable when prompts require external world knowledge. Bounded and long-tail parametric knowledge prevents direct or reason-then-generate approache…

Image GenerationImage Editing

RefDiT: Local Attribute Guidance in Reference-Based Image Generation

2026-09-04 · Rameshwar Mishra, Srikrishna Karanam, A V Subramanyam arxiv

Personalization models generate new images guided by a few subject references, while style transfer methods aim to produce images aligned with a global style derived from a reference image. Recent approaches perform well…

Image GenerationStyle Transfer

Editable Visual Design

2026-09-03 · Junyan Ye, Wei Liu, Dongzhi Jiang, Zichen Wen 외 hf

While diffusion base models such as GPT-Image-2 and Nano-Banana exhibit remarkable visual expressiveness, their end-to-end generation inherently yields flattened bitmaps with error-prone text, precluding layer-wise post-…

Image Generation

When Composition Doesn't Add Up: Humans Identifying Defects in AI-Generated Images

2026-08-26 · Ruoqi Hu, Chulin Zhao, Jiashuo Chang, Ramon Ruiz-Dolz 외 arxiv

*Chulin Zhao and Ruoqi Hu contributed equally to this work. State-of-the-art text-to-image (T2I) models exhibit pronounced and systematic defects when prompts involve intricate compositional factors such as multiple enti…

Image Generation

Efficient Training with Foresight: Multi-Token Auxiliary Supervision for Autoregressive Image Generation

2026-08-26 · Guo Niu, Xiongfei Yao, Teng Wang, Nannan Zhu arxiv

Autoregressive (AR) image generation has shown strong potential for scalable high-fidelity synthesis by modeling images as discrete token sequences. However, traditional next token prediction (NTP) continues to suffer fr…

Image Generation

ChebBooster: A Training-Free Approach for Efficient Diffusion Transformer Inference via Chebyshev-Inspired Extrapolation

2026-08-24 · Chengjie Lu, Tianchi Deng, Zhengqi He, Chengwen Luo 외 arxiv

Diffusion Transformers (DiTs) have shown strong performance in high-fidelity image generation, but their sampling process remains computationally intensive due to full model execution at every timestep. While cache-based…

Image Generation

Grounding Free-Form Instructions for Fashion Complementary Image Generation

2026-08-24 · Matteo Attimonelli, Claudio Pomo, Alessandro De Bellis, Danilo Danese 외 arxiv

Fashion complementary image generation (CIG) aims to create garments that stylistically match a seed item based on user intent, making it a natural multimodal grounding problem where models must interpret language in vis…

Image Generation

MIVIFI: Bridging Perspective and Fisheye Domains for Training Multi-View Fisheye Image Generation Models

2026-08-24 · Matthias Neuwirth-Trapp, Begüm Altunbas, Jiayi Wang, Yan Xia 외 arxiv

Achieving 360° coverage is critical for the visual perception systems of autonomous vehicles. Fisheye cameras offer a cost-effective solution by enabling full surround coverage with as few as two sensors. However, existi…

Autonomous VehiclesImage Generation

GAN-Diff : Coupling Pretrained WGAN-GP Features with Conditional Diffusion U-Nets

2026-08-23 · Saif Ahmed, Asadullah Hil Galib, S. M. Riaz Rahman Antu, Ahmed Faizul Haque Dhrubo 외 arxiv

Generative adversarial networks (GANs) can provide efficient image generation, while diffusion models offer high-quality image restoration but require iterative sampling. This paper presents a hybrid GAN-guided diffusion…

Image RestorationImage Generation

Bridging Language and Spherical Space: Object-Centric Control for Text-to-Panorama Generation

2026-08-21 · Derui Li, Qian Qiao, Yuhao Sun, Wenhao Guo 외 arxiv

Panoramic image generation is increasingly important for immersive applications such as virtual reality, augmented reality, and 3D content creation. Unlike perspective images, panoramic images represent a viewer-centered…

Spatial ReasoningImage Generation

WithEveryone: Unified Planning and Identity Grounding for Group Image Generation

2026-08-20 · Hengyuan Xu, Qixun Wang, Yiji Cheng, Miles Yang 외 arxiv

Identity-preserving image generation becomes increasingly unreliable when a scene must contain many specified people. Beyond retaining each identity, the model must bind every reference to a distinct person and location,…

Image Generation

SemanticSlider3D: Training-Free Continuous Semantic Editing for 3D Objects

2026-08-19 · Ru Wang, Rahul Jain, Koichiro Niinuma, Aakar Gupta arxiv

Fine-grained control over continuous semantic attributes of 3D objects is essential for 3D content creation, but is not well supported by conventional 3D modeling workflows or prompt-based interaction with existing gener…

Image Generation3D Generation

From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation

2026-08-18 · Xingjian Wang, Zhao Wang, Taihang Hu, Jun Zheng 외 arxiv

Large-scale image generation has benefited from advances in data scale, quality, rebalancing, and recaptioning, yet conventional pipelines typically optimize task-specific datasets in isolation. A central challenge is no…

Image Generation

TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation

2026-08-17 · Haoran Wang, Chaofan Ma, Ran Yi, Lizhuang Ma arxiv

Despite recent advances in unified multimodal models for multi-reference image generation, existing benchmarks remain organized around predefined task types (e.g., "subject composition"), which are ill-suited to this com…

Image Generation

GenRouter: Unified Workflow Routing for Agentic Image Generation

2026-08-17 · Harold Haodong Chen, Zhiyu Hou, Wen-Jie Shu, Weilin Ruan 외 arxiv

The rapid evolution of text-to-image (T2I) generation models has effectively solved the foundational challenge of raw pixel synthesis, shifting the community's focus toward fulfilling increasingly intricate user requests…

Zero-shot GeneralizationImage Generation

Defake-o3: From Speculative Rationales to Verifiable Evidence for Explainable AIGI Detection

2026-08-17 · Bowen Deng, Jiahui Zhan, Yikun Ji, Haozhen Yan 외 arxiv

The rapid progress of image generation models calls for AI-generated image (AIGI) detectors that are not only accurate but also explainable and reliable. While MLLM-based detectors can provide natural language explanatio…

Reinforcement LearningVisual GroundingImage Generation

RankT2I: A Submodular Framework for Discovering Interpretable and Diverse Semantics in Text-to-Image Models

2026-08-14 · Ritika Allada, Pinar Yanardag arxiv

Recent advances in text-to-image (T2I) models have revolutionized the field of image generation and editing. However, identifying semantics that a T2I model can successfully edit in an image continues to be a challenging…

Image GenerationImage Editing

Evaluation of Clinically Steerable Retinal Image Generation from Foundation Model Latent Spaces

2026-08-13 · Zuzanna A. Wakefield-Skórniewska, Bartłomiej W. Papież arxiv

Medical foundation models learn latent representations of clinically meaningful phenotypes, yet their ability to support controllable image generation remains largely unexplored. We evaluate four retinal foundation model…

Image Generation

From Local Mismatch to Global Impact: Optimizing Cache Reuse Policy for Efficient Diffusion

2026-08-13 · Xichen Ye, Yifan Wu, Zhikang Xie, Xiangyu Yue 외 arxiv

Diffusion models have achieved dominant performance in visual generation but suffer from substantial inference overhead. While cache-based acceleration has emerged as a promising solution, existing policies rely on local…

Bilevel OptimizationImage Generation
1–20 / 7,946 다음 →