Image Manipulation
1개 벤치마크 · 논문 486편 · 이 태스크의 논문 보기 →
Benchmarks
LRS2
Most implemented
SinGAN: Learning a Generative Model from a Single Natural Image
Closed-Form Factorization of Latent Semantics in GANs
MaskGIT: Masked Generative Image Transformer
Designing an Encoder for StyleGAN Image Manipulation
SRFlow: Learning the Super-Resolution Space with Normalizing Flow
Papers
FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification
Agentic vision-language models (VLMs), which interleave textual reasoning with explicit tool calls such as cropping and code-based image manipulation, have emerged as a compelling paradigm for reliable and interpretable …
Multimodal ReasoningImage ManipulationCan Vision-Language Models Reason about AI Edits in Images?
Detection and localization of AI-tampered images are critical for trustworthy AI, yet modern generative models have made such manipulations increasingly difficult to identify. While traditional binary classifiers can det…
Reinforcement LearningImage ManipulationEditing Everything Everywhere All at Once
Editing multiple elements of an image in a single forward pass is a practical alternative to multi-turn image manipulation, offering improved efficiency and potentially better harmonization. However, when several instruc…
Image ManipulationImage EditingEPEdit: Redefining Image Editing with Generative AI and User-Centric Design
The demand for image manipulation has seen a significant increase recently. Traditional tools like Photoshop and Capture One, while powerful, require considerable expertise to use effectively. Generative AI has introduce…
Image ManipulationImage GenerationImage EditingText-Vision Co-Instructed Image Editing
Existing image editing methods can be generally categorized into textual instruction-based and visual prompt-based ones. Textual instructions are semantically expressive, but are limited by the coarse granularity of spat…
Image ManipulationImage EditingToward 360-Degree Indoor Panorama Editing via Tuning-Free Diffusion Model with Refocusing Cross-Attention
Zero-shot text-guided diffusion has significantly advanced image editing; however, its practical usability remains constrained by three persistent challenges: prompt brittleness that requires meticulous prompt engineerin…
Image ManipulationPrompt EngineeringImage Editing