paper-with-me

Papers

Scaling Dense Event-Stream Pretraining from Visual Foundation Models

2026-03-04 · Zhiwen Chen, Junhui Hou, Zhiyu Zhu, Jinjian Wu, Guangming Shi arxiv

Learning versatile, fine-grained representations from irregular event streams is pivotal yet nontrivial, primarily due to the heavy annotation that hinders scalability in dataset size, semantic richness, and application scope. To mitigate this dilemma, we launch a novel self-supervised pretraining method that distills visual foundation models (VFMs) to push the boundaries of event representation at scale. Specifically, we curate an extensive synchronized image-event collection to amplify cross-modal alignment. Nevertheless, due to inherent mismatches in sparsity and granularity between image-event domains, existing distillation paradigms are prone to semantic collapse in event representations, particularly at high resolutions. To bridge this gap, we propose to extend the alignment objective to semantic structures provided off-the-shelf by VFMs, indicating a broader receptive field and stronger supervision. The key ingredient of our method is a structure-aware distillation loss that grounds higher-quality image-event correspondences for alignment, optimizing dense event representations. Extensive experiments demonstrate that our approach takes a great leap in downstream benchmarks, significantly surpassing traditional methods and existing pretraining techniques. This breakthrough manifests in enhanced generalization, superior data efficiency and elevated transferability.

📄 PDF Abstract BibTeX arXiv:2603.03969

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Vision Pretraining for Dense Spatial Perception

2026-07-06 · Zelin Fu, Bin Tan, Changjiang Sun, Shaohui Liu 외 arxiv

Dense spatial perception is essential for physical intelligence, where visual systems are expected to recover structured, metric, and actionable representations from pixel observations. Modern visual foundation models te…

Depth CompletionDepth Estimation

Rethinking Person Re-Identification via Semantic-Based Pretraining

2021-10-11 · Suncheng Xiang, Jingsheng Gao, Zirui Zhang, Mengyuan Guan 외

Pretraining is a dominant paradigm in computer vision. Generally, supervised ImageNet pretraining is commonly used to initialize the backbones of person re-identification (Re-ID) models. However, recent works show a surp…

Person Re-Identification

Generative Medical Event Models Improve with Scale

2025-08-16 · Shane Waxler, Paul Blazek, Davis White, Daniel Sneider 외 arxiv

Realizing personalized medicine at scale calls for methods that distill insights from longitudinal patient journeys, which can be viewed as a sequence of medical events. Foundation models pretrained on large-scale medica…

Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers

2026-07-30 · Chongjian Ge, Hanwen Jiang, Tianyu Wang, Jiuxiang Gu 외 arxiv

Visual generation increasingly requires high-resolution images, long videos, and multimodal context, making the quadratic cost of full attention prohibitive. We introduce Chimera, a hybrid visual diffusion backbone with …

Beyond Language Modeling: An Exploration of Multimodal Pretraining

2026-03-03 · Shengbang Tong, David Fan, John Nguyen, Ellis Brown 외 arxiv

The visual world offers a critical axis for advancing foundation models beyond language. Despite growing interest in this direction, the design space for native multimodal models remains opaque. We provide empirical clar…