paper-with-me

홈 › Papers

From Sky to the Ground: A Large-scale Benchmark and Simple Baseline Towards Real Rain Removal

2023-08-07 · ICCV 2023 1 · Yun Guo, Xueyao Xiao, Yi Chang, Shumin Deng, Luxin Yan

Learning-based image deraining methods have made great progress. However, the lack of large-scale high-quality paired training samples is the main bottleneck to hamper the real image deraining (RID). To address this dilemma and advance RID, we construct a Large-scale High-quality Paired real rain benchmark (LHP-Rain), including 3000 video sequences with 1 million high-resolution (1920*1080) frame pairs. The advantages of the proposed dataset over the existing ones are three-fold: rain with higher-diversity and larger-scale, image with higher-resolution and higher-quality ground-truth. Specifically, the real rains in LHP-Rain not only contain the classical rain streak/veiling/occlusion in the sky, but also the \textbf{splashing on the ground} overlooked by deraining community. Moreover, we propose a novel robust low-rank tensor recovery model to generate the GT with better separating the static background from the dynamic rain. In addition, we design a simple transformer-based single image deraining baseline, which simultaneously utilize the self-attention and cross-layer attention within the image and rain layer with discriminative feature representation. Extensive experiments verify the superiority of the proposed dataset and deraining method over state-of-the-art.

📄 PDF Abstract BibTeX arXiv:2308.03867

Code (1)

yunguo224/lhp-rain 공식 구현 pytorch

Tasks

DiversityRain RemovalSingle Image Deraining

Similar Papers 제목 키워드 기반

When Visual Grounding Meets Gigapixel-level Large-scale Scenes: Benchmark and Approach

2024-01-01 · CVPR 2024 1 · Tao Ma, Bing Bai, Haozhe Lin, Heyuan Wang 외

Visual grounding refers to the process of associating natural language expressions with corresponding regions within an image. Existing benchmarks for visual grounding primarily operate within small-scale scenes with…

Scene UnderstandingVisual Grounding

SimBase: A Simple Baseline for Temporal Video Grounding

2024-11-12 · Peijun Bao, Alex C. Kot

This paper presents SimBase, a simple yet effective baseline for temporal video grounding. While recent advances in temporal grounding have led to impressive performance, they have also driven network architectures towar…

Video Grounding

PixFoundation: Are We Heading in the Right Direction with Pixel-level Vision Foundation Models?

2025-02-06 · Mennatullah Siam

Multiple works have emerged to push the boundaries on multi-modal large language models (MLLMs) towards pixel-level understanding. Such approaches have shown strong performance on benchmarks for referring expression segm…

Question AnsweringReferring ExpressionReferring Expression SegmentationVisual Question Answering

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation

2026-04-17 · Shuyan Ke, Yifan Mei, Changli Wu, Yonghan Zheng 외 arxiv

Reasoning segmentation has recently expanded from ground-level scenes to remote-sensing imagery, yet UAV data poses distinct challenges, including oblique viewpoints, ultra-high resolutions, and extreme scale variations.…

Have We Mastered Scale in Deep Monocular Visual SLAM? The ScaleMaster Dataset and Benchmark

2026-02-20 · Hyoseok Ju, Bokeon Suh, Giseop Kim arxiv

Recent advances in deep monocular visual Simultaneous Localization and Mapping (SLAM) have achieved impressive accuracy and dense reconstruction capabilities, yet their robustness to scale inconsistency in large-scale in…