paper-with-me

홈 › Papers

Artificial-Spiking Hierarchical Networks for Vision-Language Representation Learning

2023-08-18 · Yeming Chen, Siyu Zhang, Yaoru Sun, Weijian Liang, Haoran Wang

With the success of self-supervised learning, multimodal foundation models have rapidly adapted a wide range of downstream tasks driven by vision and language (VL) pretraining. State-of-the-art methods achieve impressive performance by pre-training on large-scale datasets. However, bridging the semantic gap between the two modalities remains a nonnegligible challenge for VL tasks. In this work, we propose an efficient computation framework for multimodal alignment by introducing a novel visual semantic module to further improve the performance of the VL tasks. Specifically, we propose a flexible model, namely Artificial-Spiking Hierarchical Networks (ASH-Nets), which combines the complementary advantages of Artificial neural networks (ANNs) and Spiking neural networks (SNNs) to enrich visual semantic representations. In particular, a visual concrete encoder and a semantic abstract encoder are constructed to learn continuous and discrete latent variables to enhance the flexibility of semantic encoding. Considering the spatio-temporal properties of SNNs modeling, we introduce a contrastive learning method to optimize the inputs of similar samples. This can improve the computational efficiency of the hierarchical network, while the augmentation of hard samples is beneficial to the learning of visual representations. Furthermore, the Spiking to Text Uni-Alignment Learning (STUA) pre-training method is proposed, which only relies on text features to enhance the encoding ability of abstract semantics. We validate the performance on multiple well-established downstream VL tasks. Experiments show that the proposed ASH-Nets achieve competitive results.

📄 PDF Abstract BibTeX arXiv:2308.09455

Code (0)

등록된 구현이 없습니다.

Tasks

Computational EfficiencyContrastive LearningRepresentation LearningSelf-Supervised Learning

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

SpikingSSMs: Learning Long Sequences with Sparse and Parallel Spiking State Space Models

2024-08-27 · Shuaijie Shen, Chao Wang, Renzhuo Huang, Yan Zhong 외

Known as low energy consumption networks, spiking neural networks (SNNs) have gained a lot of attention within the past decades. While SNNs are increasing competitive with artificial neural networks (ANNs) for vision tas…

Language ModelingLanguage ModellingState Space Models

QKFormer: Hierarchical Spiking Transformer using Q-K Attention

2024-03-25 · Chenlin Zhou, Han Zhang, Zhaokun Zhou, Liutao Yu 외

Spiking Transformers, which integrate Spiking Neural Networks (SNNs) with Transformer architectures, have attracted significant attention due to their potential for energy efficiency and high performance. However, existi…

Winner-Take-All Spiking Transformer for Language Modeling

2026-04-13 · Chenlin Zhou, Sihang Guo, Jiaqi Wang, Dongyang Ma 외 arxiv

Spiking Transformers, which combine the scalability of Transformers with the sparse, energy-efficient property of Spiking Neural Networks (SNNs), have achieved impressive results in neuromorphic and vision tasks and attr…

Natural Language Understanding

The Spike Gating Flow: A Hierarchical Structure Based Spiking Neural Network for Online Gesture Recognition

2022-06-04 · Zihao Zhao, Yanhong Wang, Qiaosha Zou, Tie XU 외

Action recognition is an exciting research avenue for artificial intelligence since it may be a game changer in the emerging industrial fields such as robotic visions and automobiles. However, current deep learning faces…

Action RecognitionFew-Shot LearningGesture Recognition

Exploring the Potentials of Spiking Neural Networks for Image Deraining

2025-12-01 · Shuang Chen, Tomas Krajnik, Farshad Arvin, Amir Atapour-Abarghouei arxiv

Biologically plausible and energy-efficient frameworks such as Spiking Neural Networks (SNNs) have not been sufficiently explored in low-level vision tasks. Taking image deraining as an example, this study addresses the …

Representation Learning