paper-with-me

Papers Image Editing

“Image Editing” 태그가 달린 논문 608편 · 필터 해제

Mi-Ripple: Restoring Images Degraded by Iterative AI Editing

2026-09-10 · Jiayin Chen, Yicheng Xu, Muting Wang hf

Iterative reference-conditioned image editing can introduce grid-like and granular textures, commonly described as digital ripple. We present Mi-Ripple, a diagnosis-guided restoration workflow that suppresses this digita…

Image Editing

SenseNova-U1.5: Towards Native Unified Visual Intelligence

2026-09-10 · Haiwen Diao, Jiahao Wang, Chenjing Ding, Hanming Deng 외 hf

We launch SenseNova-U1.5, an 8B-MoT native unified multimodal model that understands, reasons about, and generates visual content within an encoder-free and VAE-free architecture. We strengthen its visual interface throu…

Reinforcement LearningInstruction FollowingImage Editing

WeAgent-MMGenEdit: A Full-Stack Recipe for Multimodal Agentic Image Generation and Editing

2026-09-04 · Hui Zhang, Zongkai Liu, Liqiang Niu, Juntao Liu 외 arxiv

Image generation and editing models have advanced rapidly, yet remain unreliable when prompts require external world knowledge. Bounded and long-tail parametric knowledge prevents direct or reason-then-generate approache…

Image GenerationImage Editing

Learning 3D Editing without Paired Supervision via Generative Prior Distillation

2026-09-04 · Hao Wen, Weibin Yun, Hongxing Fan, Haotian Lu 외 arxiv

Instruction-guided 3D editing is essential for interactive content creation, yet it faces a significant bottleneck: the severe scarcity of high-quality paired training data. Existing approaches attempt to bypass this by …

Instruction FollowingImage Editing

LensStyle: Learning the Optical Aesthetics for Controllable Stylized Lens Effect Rendering

2026-09-04 · Yachuan Huang, Liwen Xiao, Liao Shen, Qiwen Wang 외 arxiv

The visual aesthetics of photographs are deeply influenced by lens characteristics such as aperture shape, optical vignetting and optical diffraction, which together define a camera's unique optical style. Existing lens …

Image Editing

AngelFingerprint: A Traceable, Explainable, and White-Box Stealthy Watermark for Text-Guided Image Editing

2026-09-04 · Bo-Han Kung, Futa Waseda, Ching-Chun Chang, Isao Echizen 외 arxiv

Text-guided diffusion editing raises disinformation concerns, making reliable image provenance essential. While watermarks are commonly used for this purpose, most methods carry a fixed ID that cannot explain what was ch…

Image Editing

An Evaluation Framework for Generating Multi-View Images of a Person in a Scene

2026-09-04 · Mahir Majid, Young Kyung Kim, Guillermo Sapiro arxiv

Recent generative image-editing Diffusion Transformers (DiTs) demonstrate impressive semantic editing capabilities but still struggle with spatially consistent camera angle changes. A primary bottleneck in training found…

Image Editing

Aphanta: Diagnosing Task-Aligned Image-Edited Intermediates for Multimodal Reasoning

2026-08-27 · Hengyuan Xu, Wei Cheng, Yumeng Ji, Xuanyang Zhang 외 hf

Explicit visual intermediates can help multimodal large language models (MLLMs) externalize spatial evidence and updated visual states, but their utility depends on whether an image editor can faithfully realize the requ…

Multimodal ReasoningImage Editing

RubricRM: Generative Reward Modeling via Dynamic Rubrics for Image Generation and Editing

2026-08-27 · Zijian Kan, Wei Wang, Long Luo, Bing Zhao 외 arxiv

Reward models play an essential role in aligning visual generative models, yet most existing visual reward models use a single scalar score or rely on fixed criteria that cannot adapt to different instructions. This limi…

Text-to-Image GenerationImage Editing

RefVideo-6M: A Reliable Reference-Based Dataset for Instructional Video Editing

2026-08-26 · Bojia Zi, Xiaoyan Yang, Yu Zhou, Ruijie Sun 외 arxiv

Recent advances in video editing have been largely driven by large-scale instruction-based datasets. However, existing datasets still suffer from two critical limitations. First, target videos are commonly produced by au…

Image Editing

FlashNormal: Detailed Surface Normal Estimation from Flash and No-Flash Images

2026-08-26 · Ruiyang Chen, Feiran Li, Heng Guo, Zhanyu Ma arxiv

High-quality surface normal estimation is preferred for detailed surface shape recovery and image editing. Existing single image-based methods, though being a practical setup, often struggle to recover fine surface detai…

Image Editing

When the Edit Changes the Patient: Measuring Identity Preservation in Counterfactual Retinal Images

2026-08-24 · Andrea Posada, Wenke Karbole, Bach Ngoc Doan, Alexander Weers 외 arxiv

Counterfactual medical image generation aims to modify an existing image to reflect a hypothetical scenario in which certain characteristics of the imaged subject are altered, while keeping their identity fixed. Most exi…

Medical Image GenerationImage Editing

Can We Perform Online RL for Image Editing without Editing Rewards?

2026-08-24 · Qichao Ma, Jikang Cheng, Ling Liang, Zhaofei Yu 외 arxiv

Reinforcement learning (RL) enables direct preference optimization for image editing through editing-specific rewards, which remain less developed due to costly triplet supervision and complex task-dependent calibration.…

Reinforcement LearningImage Editing

Exploring the Performance Frontier of Compact Unified Image Generation Models

2026-08-20 · Taihang Hu, Zhao Wang, Zuan Gao, Tao Liu 외 arxiv

We present Swift-Image, a compact unified model for text-to-image generation, single-image editing, and multi-image editing. Our goal is to explore how far a relatively small visual generator can be pushed through system…

Text-to-Image GenerationReinforcement LearningModel CompressionImage Editing

DARS: Dual-Level Credit Assignment RL with Structured Reasoning for Instruction-Based Image Editing

2026-08-20 · Haoxiang Cao, Jiajiong Cao, Xuanpu Zhang, Changqian Yu 외 arxiv

Instruction-based image editing uses a planner-renderer pipeline: a vision-language model (VLM) first converts the instruction into an edit plan, and a diffusion model then executes that plan. Training such systems with …

Reinforcement LearningImage Editing

TextRefine: Improving Textual Fidelity, Spatial Placement, and Glyph Rendering for Text Editing in Product Posters

2026-08-20 · Honglie Wang, Jia Sun, Zijun Li, Junlong Wu 외 arxiv

Text editing in product posters entails inserting new text or replacing existing text while preserving product appearance, background content, and global composition. Despite recent progress in instruction-based image ed…

Image Editing

EDITBRIDGE: Towards Faithful and Efficient Ultra-High-Resolution Image Editing

2026-08-18 · Jiayi Song, Shijie Huang, Fangtai Wu, Yubo Huang 외 arxiv

High-resolution image editing is increasingly demanded in professional workflows, yet existing diffusion-based models remain constrained to resolutions below 1K due to quadratic attention complexity and prohibitive memor…

Semantic correspondenceImage Editing

Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision

2026-08-17 · Long Cui, Xiaoqian Liu, Qi Qin, Yi Xin 외 arxiv

Existing image editing frameworks predominantly follow the training paradigm of text-to-image diffusion models. However, extending this paradigm to image editing highlights two inherent discrepancies, specifically, the i…

Image Editing

TransAnyText: Translating Arbitrary Text in E-commerce Images via Structured Visual Generation

2026-08-17 · Xiaoan Liu, Lichen Ma, Zipeng Guo, Yu He 외 arxiv

Cross-border e-commerce image translation is essential for global retail, where product images, banners, and detail pages need to be produced in different languages. Existing methods struggle to achieve accurate translat…

Reinforcement LearningImage Editing

CPI-Bench: A Comprehensive,Practical and Intelligent Benchmark for Real-World Image Editing

2026-08-14 · Qinye Zhou, Jun Zheng, Yongchao Du, Yuan Wang 외 arxiv

With the rapid advancement of image editing models and their widespread application across various domains, there is an increasingly urgent need to deploy these model capabilities directly into real-world scenarios. Howe…

Image Editing
1–20 / 608 다음 →