paper-with-me

홈 › Papers

UniProcessor: A Text-induced Unified Low-level Image Processor

2024-07-30 · Huiyu Duan, Xiongkuo Min, Sijing Wu, Wei Shen, Guangtao Zhai

Image processing, including image restoration, image enhancement, etc., involves generating a high-quality clean image from a degraded input. Deep learning-based methods have shown superior performance for various image processing tasks in terms of single-task conditions. However, they require to train separate models for different degradations and levels, which limits the generalization abilities of these models and restricts their applications in real-world. In this paper, we propose a text-induced unified image processor for low-level vision tasks, termed UniProcessor, which can effectively process various degradation types and levels, and support multimodal control. Specifically, our UniProcessor encodes degradation-specific information with the subject prompt and process degradations with the manipulation prompt. These context control features are injected into the UniProcessor backbone via cross-attention to control the processing procedure. For automatic subject-prompt generation, we further build a vision-language model for general-purpose low-level degradation perception via instruction tuning techniques. Our UniProcessor covers 30 degradation types, and extensive experiments demonstrate that our UniProcessor can well process these degradations without additional training or tuning and outperforms other competing methods. Moreover, with the help of degradation-aware context control, our UniProcessor first shows the ability to individually handle a single distortion in an image with multiple degradations.

📄 PDF Abstract BibTeX arXiv:2407.20928

Code (1)

intmegroup/uniprocessor 공식 구현 pytorch

Tasks

Image EnhancementImage RestorationLanguage Modelling

Similar Papers 제목 키워드 기반

Fixed-Priority and EDF Schedules for ROS2 Graphs on Uniprocessor

2025-11-28 · Oren Bell, Harun Teper, Mario Günzel, Chris Gill 외 arxiv

This paper addresses limitations of current scheduling methods in the Robot Operating System (ROS)2, focusing on scheduling tasks beyond simple chains and analyzing arbitrary Directed Acyclic Graphs (DAGs). While previou…

Can We Build a Monolithic Model for Fake Image Detection? SICA: Semantic-Induced Constrained Adaptation for Unified-Yet-Discriminative Artifact Feature Space Reconstruction

2026-02-06 · Bo Du, Xiaochen Ma, Xuekang Zhu, Zhe Yang 외 arxiv

Fake Image Detection (FID), aiming at unified detection across four image forensic subdomains, is critical in real-world forensic scenarios. Compared with ensemble approaches, monolithic FID models are theoretically more…

Train a Unified Multimodal Data Quality Classifier with Synthetic Data

2025-10-16 · Weizhi Wang, Rongmei Lin, Shiyang Li, Colin Lockard 외 arxiv

The Multimodal Large Language Models (MLLMs) are continually pre-trained on a mixture of image-text caption data and interleaved document data, while the high-quality data filtering towards image-text interleaved documen…

Generalized Decoding for Pixel, Image, and Language

2022-12-21 · CVPR 2023 1 · Xueyan Zou, Zi-Yi Dou, Jianwei Yang, Zhe Gan 외

We present X-Decoder, a generalized decoding model that can predict pixel-level segmentation and language tokens seamlessly. X-Decodert takes as input two types of queries: (i) generic non-semantic queries and (ii) seman…

DecoderImage SegmentationInstance SegmentationPanoptic Segmentation+4

SkyNative: A Native Multimodal Framework for Remote Sensing Visual Evidence Reasoning

2026-05-18 · Xiao Yang, Ronghao Fu, Zhiwen Lin, Zhuoran Duan 외 arxiv

Remote sensing vision-language models commonly rely on pretrained visual encoders to convert images into semantic features before language-model reasoning. While effective for scene-level understanding, this pipeline may…

Spatial Reasoning