Papers Image Manipulation
“Image Manipulation” 태그가 달린 논문 486편 · 필터 해제
FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification
Agentic vision-language models (VLMs), which interleave textual reasoning with explicit tool calls such as cropping and code-based image manipulation, have emerged as a compelling paradigm for reliable and interpretable …
Multimodal ReasoningImage ManipulationCan Vision-Language Models Reason about AI Edits in Images?
Detection and localization of AI-tampered images are critical for trustworthy AI, yet modern generative models have made such manipulations increasingly difficult to identify. While traditional binary classifiers can det…
Reinforcement LearningImage ManipulationEditing Everything Everywhere All at Once
Editing multiple elements of an image in a single forward pass is a practical alternative to multi-turn image manipulation, offering improved efficiency and potentially better harmonization. However, when several instruc…
Image ManipulationImage EditingEPEdit: Redefining Image Editing with Generative AI and User-Centric Design
The demand for image manipulation has seen a significant increase recently. Traditional tools like Photoshop and Capture One, while powerful, require considerable expertise to use effectively. Generative AI has introduce…
Image ManipulationImage GenerationImage EditingText-Vision Co-Instructed Image Editing
Existing image editing methods can be generally categorized into textual instruction-based and visual prompt-based ones. Textual instructions are semantically expressive, but are limited by the coarse granularity of spat…
Image ManipulationImage EditingToward 360-Degree Indoor Panorama Editing via Tuning-Free Diffusion Model with Refocusing Cross-Attention
Zero-shot text-guided diffusion has significantly advanced image editing; however, its practical usability remains constrained by three persistent challenges: prompt brittleness that requires meticulous prompt engineerin…
Image ManipulationPrompt EngineeringImage EditingComparative Evaluation of Deep Learning Models for Fake Image Detection
The growing sophistication of GAN-based image manipulation presents significant challenges for digital forensics. This study compares the performance of four pretrained CNN architectures including VGG16, ResNet50, Effici…
Image ManipulationAre Watermarked Images Editable? SafeMark for Watermark-Preserving Text-Guided Image Editing
This paper investigates a fundamental yet underexplored question: can watermarked images remain editable without compromising watermark integrity? We propose SafeMark, a framework for watermark-preserving text-guided ima…
Image ManipulationImage EditingVenus-DeFakerOne: Unified Fake Image Detection & Localization
In recent years, the rapid evolution of generative AI has fundamentally reshaped the paradigm of image forgery, breaking the traditional boundaries between document editing, natural image manipulation, DeepFake generatio…
Image ManipulationFeatMap: Understanding image manipulation in the feature space and its implications for feature space geometry
Intermediate feature representations represent the backbone for the expressivity and adaptability of deep neural networks. However, their geometric structure remains poorly understood. In this submission, we provide indi…
Image ManipulationImage EditingEditSleuth: A Dataset of Grounded Reasoning Chains for Image-Edit Forensics
Forensic analysis of AI-edited images requires more than binary real-versus-fake prediction: a useful system should localize the edit, identify its semantic type, and ground its decisions in visual evidence. Existing ima…
Image ManipulationLIVE: Leveraging Image Manipulation Priors for Instruction-based Video Editing
Video editing aims to modify input videos according to user intent. Recently, end-to-end training methods have garnered widespread attention, constructing paired video editing data through video generation or editing mod…
Image ManipulationVideo GenerationImage EditingAIM-Bench: Benchmarking and Improving Affective Image Manipulation via Fine-Grained Hierarchical Control
Affective Image Manipulation (AIM) aims to evoke specific emotions through targeted editing. Current image editing benchmarks primarily focus on object-level modifications in general scenarios, lacking the fine-grained g…
Image ManipulationImage EditingImmune2V: Image Immunization Against Dual-Stream Image-to-Video Generation
Image-to-video (I2V) generation has the potential for societal harm because it enables the unauthorized animation of static images to create realistic deepfakes. While existing defenses effectively protect against static…
Image ManipulationVideo GenerationSeeing the Evidence, Missing the Answer: Tool-Guided Vision-Language Models on Visual Illusions
Vision-language models (VLMs) exhibit a systematic bias when confronted with classic optical illusions: they overwhelmingly predict the illusion as "real" regardless of whether the image has been counterfactually modifie…
Image ManipulationSpatial ReasoningImage CompressionCREval: An Automated Interpretable Evaluation for Creative Image Manipulation under Complex Instructions
Instruction-based multimodal image manipulation has recently made rapid progress. However, existing evaluation methods lack a systematic and human-aligned framework for assessing model performance on complex and creative…
Image ManipulationImage EditingAdaEdit: Adaptive Temporal and Channel Modulation for Flow-Based Image Editing
Inversion-based image editing in flow matching models has emerged as a powerful paradigm for training-free, text-guided image manipulation. A central challenge in this paradigm is the injection dilemma: injecting source …
Image ManipulationImage EditingVisual Prompt Discovery via Semantic Exploration
LVLMs encounter significant challenges in image understanding and visual reasoning, leading to critical perception failures. Visual prompts, which incorporate image manipulation code, have shown promising potential in mi…
Image ManipulationVisual ReasoningEdit2Interp: Adapting Image Foundation Models from Spatial Editing to Video Frame Interpolation with Few-Shot Learning
Pre-trained image editing models exhibit strong spatial reasoning and object-aware transformation capabilities acquired from billions of image-text pairs, yet they possess no explicit temporal modeling. This paper demons…
Video Frame InterpolationImage ManipulationSpatial ReasoningFew-Shot LearningRethinking VLMs for Image Forgery Detection and Localization
With the rapid rise of Artificial Intelligence Generated Content (AIGC), image manipulation has become increasingly accessible, posing significant challenges for image forgery detection and localization (IFDL). In this p…
Image Manipulation