paper-with-me

Papers

SPKLIP: Aligning Spike Video Streams with Natural Language

2025-05-19 · Yongchang Gao, Meiling Jin, Zhaofei Yu, Tiejun Huang, Guozhang Chen

Spike cameras offer unique sensing capabilities but their sparse, asynchronous output challenges semantic understanding, especially for Spike Video-Language Alignment (Spike-VLA) where models like CLIP underperform due to modality mismatch. We introduce SPKLIP, the first architecture specifically for Spike-VLA. SPKLIP employs a hierarchical spike feature extractor that adaptively models multi-scale temporal dynamics in event streams, and uses spike-text contrastive learning to directly align spike video with language, enabling effective few-shot learning. A full-spiking visual encoder variant, integrating SNN components into our pipeline, demonstrates enhanced energy efficiency. Experiments show state-of-the-art performance on benchmark spike datasets and strong few-shot generalization on a newly contributed real-world dataset. SPKLIP's energy efficiency highlights its potential for neuromorphic deployment, advancing event-based multimodal research. The source code and dataset are available at [link removed for anonymity].

📄 PDF Abstract BibTeX arXiv:2505.12656

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningFew-Shot Learning

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음
SNN Spiking Neural Networks (SNNs) are a class of artificial neural networks inspired by the structure and functioning of the brain's neural networks. Unlike traditional…
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Spike-guided Motion Deblurring with Unknown Modal Spatiotemporal Alignment

2024-01-01 · CVPR 2024 1 · Jiyuan Zhang, Shiyan Chen, Yajing Zheng, Zhaofei Yu 외

The traditional frame-based cameras that rely on exposure windows for imaging experience motion blur in high-speed scenarios. Frame-based deblurring methods lack reliable motion cues to restore sharp images under ext…

DeblurringImage Deblurring

SpikeGen: Generative Framework for Visual Spike Stream Processing

2025-05-23 · Gaole Dai, Menghang Dong, Rongyu Zhang, Ruichuan An 외

Neuromorphic Visual Systems, such as spike cameras, have attracted considerable attention due to their ability to capture clear textures under dynamic conditions. This capability effectively mitigates issues related to m…

DeblurringNovel View SynthesisVideo Deblurring

SpikeDerain: Unveiling Clear Videos from Rainy Sequences Using Color Spike Streams

2025-03-26 · Hanwen Liang, Xian Zhong, Wenxuan Liu, Yajing Zheng 외

Restoring clear frames from rainy videos presents a significant challenge due to the rapid motion of rain streaks. Traditional frame-based visual sensors, which capture scene content synchronously, struggle to capture th…

Rain RemovalVideo deraining

SEDformer: Event-Synchronous Spiking Transformers for Irregular Telemetry Time Series Forecasting

2026-02-02 · Ziyu Zhou, Yuchen Fang, Weilin Ruan, Shiyu Wang 외 arxiv

Telemetry streams from large-scale Internet-connected systems (e.g., IoT deployments and online platforms) naturally form an irregular multivariate time series (IMTS) whose accurate forecasting is operationally vital. A …

Time Series Forecasting

Seeing the Unseen in Low-light Spike Streams

2025-09-27 · Liwen Hu, Yang Li, Mianzhi Liu, Yijia Guo 외 arxiv

Spike camera, a type of neuromorphic sensor with high-temporal resolution, shows great promise for high-speed visual tasks. Unlike traditional cameras, spike camera continuously accumulates photons and fires asynchronous…