paper-with-me

Papers

Temporal-Guided Visual Foundation Models for Event-Based Vision

2025-11-09 · Ruihao Xia, Junhong Cai, Luziwei Leng, Liuyi Wang, Chengju Liu, Ran Cheng, Yang Tang, Pan Zhou arxiv

Event cameras offer unique advantages for vision tasks in challenging environments, yet processing asynchronous event streams remains an open challenge. While existing methods rely on specialized architectures or resource-intensive training, the potential of leveraging modern Visual Foundation Models (VFMs) pretrained on image data remains under-explored for event-based vision. To address this, we propose Temporal-Guided VFM (TGVFM), a novel framework that integrates VFMs with our temporal context fusion block seamlessly to bridge this gap. Our temporal block introduces three key components: (1) Long-Range Temporal Attention to model global temporal dependencies, (2) Dual Spatiotemporal Attention for multi-scale frame correlation, and (3) Deep Feature Guidance Mechanism to fuse semantic-temporal features. By retraining event-to-video models on real-world data and leveraging transformer-based VFMs, TGVFM preserves spatiotemporal dynamics while harnessing pretrained representations. Experiments demonstrate SoTA performance across semantic segmentation, depth estimation, and object detection, with improvements of 16%, 21%, and 16% over existing methods, respectively. Overall, this work unlocks the cross-modality potential of image-based VFMs for event-based vision with temporal reasoning. Code is available at https://github.com/XiaRho/TGVFM.

📄 PDF Abstract BibTeX arXiv:2511.06238

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic SegmentationEvent-based visionDepth EstimationObject Detection

Similar Papers 제목 키워드 기반

VISTA: Validation-Guided Integration of Spatial and Temporal Foundation Models with Anatomical Decoding for Rare-Pathology VCE Event Detection -- after competition results

2026-05-21 · Bo-Cheng Qiu, Fang-Ying Lin, Ming-Han Sun, Yu-Fan Lin 외 arxiv

Capsule endoscopy event detection is challenging because clinically relevant findings are sparse, visually heterogeneous, and evaluated at the event level rather than by frame accuracy. We propose VISTA, a metric-aligned…

VISTA: Validation-Guided Integration of Spatial and Temporal Foundation Models with Anatomical Decoding for Rare-Pathology VCE Event Detection

2026-03-18 · Bo-Cheng Qiu, Yu-Fan Lin, Yu-Zhe Pien, Chia-Ming Lee 외 arxiv

Capsule endoscopy event detection is challenging because diagnostically relevant findings are sparse, visually heterogeneous, and embedded in long, noisy video streams, while evaluation is performed at the event level ra…

Generative Event Pretraining with Foundation Model Alignment

2026-03-24 · Jianwen Cao, Jiaxu Xing, Nico Messikommer, Davide Scaramuzza arxiv

Event cameras provide robust visual signals under fast motion and challenging illumination conditions thanks to their microsecond latency and high dynamic range. However, their unique sensing characteristics and limited …

Object RecognitionDepth Estimation

ControlEvents: Controllable Synthesis of Event Camera Datawith Foundational Prior from Image Diffusion Models

2025-09-26 · Yixuan Hu, Yuxuan Xue, Simon Klenk, Daniel Cremers 외 arxiv

In recent years, event cameras have gained significant attention due to their bio-inspired properties, such as high temporal resolution and high dynamic range. However, obtaining large-scale labeled ground-truth data for…

Event-based visionPose Estimation

Adapting Depth Anything to Adverse Imaging Conditions with Events

2026-01-05 · Shihan Peng, Yuyang Xiong, Hanyu Zhou, Zhiwei Shi 외 arxiv

Robust depth estimation under dynamic and adverse lighting conditions is essential for robotic systems. Currently, depth foundation models, such as Depth Anything, achieve great success in ideal scenes but remain challen…

Depth Estimation