paper-with-me

Papers

Rethinking High-speed Image Reconstruction Framework with Spike Camera

2025-01-08 · Kang Chen, Yajing Zheng, Tiejun Huang, Zhaofei Yu

Spike cameras, as innovative neuromorphic devices, generate continuous spike streams to capture high-speed scenes with lower bandwidth and higher dynamic range than traditional RGB cameras. However, reconstructing high-quality images from the spike input under low-light conditions remains challenging. Conventional learning-based methods often rely on the synthetic dataset as the supervision for training. Still, these approaches falter when dealing with noisy spikes fired under the low-light environment, leading to further performance degradation in the real-world dataset. This phenomenon is primarily due to inadequate noise modelling and the domain gap between synthetic and real datasets, resulting in recovered images with unclear textures, excessive noise, and diminished brightness. To address these challenges, we introduce a novel spike-to-image reconstruction framework SpikeCLIP that goes beyond traditional training paradigms. Leveraging the CLIP model's powerful capability to align text and images, we incorporate the textual description of the captured scene and unpaired high-quality datasets as the supervision. Our experiments on real-world low-light datasets U-CALTECH and U-CIFAR demonstrate that SpikeCLIP significantly enhances texture details and the luminance balance of recovered images. Furthermore, the reconstructed images are well-aligned with the broader visual features needed for downstream tasks, ensuring more robust and versatile performance in challenging environments.

📄 PDF Abstract BibTeX arXiv:2501.04477

Code (1)

chenkang455/spikeclip 공식 구현 pytorch

Tasks

Image Reconstruction

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Rethinking Medical Image Reconstruction via Shape Prior, Going Deeper and Faster: Deep Joint Indirect Registration and Reconstruction

2019-12-16 · Jiulong Liu, Angelica I. Aviles-Rivero, Hui Ji, Carola-Bibiane Schönlieb

Indirect image registration is a promising technique to improve image reconstruction quality by providing a shape prior for the reconstruction task. In this paper, we propose a novel hybrid method that seeks to reconstru…

Computed Tomography (CT)Image ReconstructionImage Registration

Ultraman: Single Image 3D Human Reconstruction with Ultra Speed and Detail

2024-03-18 · Mingjin Chen, JunHao Chen, Xiaojun Ye, Huan-ang Gao 외

3D human body reconstruction has been a challenge in the field of computer vision. Previous methods are often time-consuming and difficult to capture the detailed appearance of the human body. In this paper, we propose a…

Lifelike 3D Human Generation

Rethinking Reconstruction and Denoising in the Dark: New Perspective, General Architecture and Beyond

2025-01-01 · CVPR 2025 1 · Tengyu Ma, Long Ma, Ziye Li, Yuetong Wang 외

Recently, enhancing image quality in the original RAW domain has garnered significant attention, with denoising and reconstruction emerging as fundamental tasks. Although some works attempt to couple these tasks, the…

Denoising

PixMIM: Rethinking Pixel Reconstruction in Masked Image Modeling

2023-03-04 · YuAn Liu, Songyang Zhang, Jiacheng Chen, Kai Chen 외

Masked Image Modeling (MIM) has achieved promising progress with the advent of Masked Autoencoders (MAE) and BEiT. However, subsequent works have complicated the framework with new auxiliary tasks or extra pre-trained mo…

Self-Supervised Learning

Rethinking Rainy 3D Scene Reconstruction via Perspective Transforming and Brightness Tuning

2025-11-10 · Qianfeng Yang, Xiang Chen, Pengpeng Li, Qiyuan Guan 외 arxiv

Rain degrades the visual quality of multi-view images, which are essential for 3D scene reconstruction, resulting in inaccurate and incomplete reconstruction results. Existing datasets often overlook two critical charact…