paper-with-me

홈 › Papers

See What You See: Self-supervised Cross-modal Retrieval of Visual Stimuli from Brain Activity

2022-08-07 · Zesheng Ye, Lina Yao, Yu Zhang, Sylvia Gustin

Recent studies demonstrate the use of a two-stage supervised framework to generate images that depict human perception to visual stimuli from EEG, referring to EEG-visual reconstruction. They are, however, unable to reproduce the exact visual stimulus, since it is the human-specified annotation of images, not their data, that determines what the synthesized images are. Moreover, synthesized images often suffer from noisy EEG encodings and unstable training of generative models, making them hard to recognize. Instead, we present a single-stage EEG-visual retrieval paradigm where data of two modalities are correlated, as opposed to their annotations, allowing us to recover the exact visual stimulus for an EEG clip. We maximize the mutual information between the EEG encoding and associated visual stimulus through optimization of a contrastive self-supervised objective, leading to two additional benefits. One, it enables EEG encodings to handle visual classes beyond seen ones during training, since learning is not directed at class annotations. In addition, the model is no longer required to generate every detail of the visual stimulus, but rather focuses on cross-modal alignment and retrieves images at the instance level, ensuring distinguishable model output. Empirical studies are conducted on the largest single-subject EEG dataset that measures brain activities evoked by image stimuli. We demonstrate the proposed approach completes an instance-level EEG-visual retrieval task which existing methods cannot. We also examine the implications of a range of EEG and visual encoder structures. Furthermore, for a mostly studied semantic-level EEG-visual classification task, despite not using class annotations, the proposed method outperforms state-of-the-art supervised EEG-visual reconstruction approaches, particularly on the capability of open class recognition.

📄 PDF Abstract BibTeX arXiv:2208.03666

Code (0)

등록된 구현이 없습니다.

Tasks

cross-modal alignmentCross-Modal RetrievalEEGElectroencephalogram (EEG)Retrieval

Similar Papers 제목 키워드 기반

How AI Experiences Art: Emergent Aesthetic Structure in a Self-Supervised Multimodal Embedding Space

2026-08-27 · Corey D. C. Heath arxiv

Aesthetics are an important part of the symbolism of artistic works. Although subjective, humans categorize art based on the emotion evoked regardless of modality. What remains under-explored is how AI models form their …

Exploiting Transformation Invariance and Equivariance for Self-supervised Sound Localisation

2022-06-26 · Jinxiang Liu, Chen Ju, Weidi Xie, Ya zhang

We present a simple yet effective self-supervised framework for audio-visual representation learning, to localize the sound source in videos. To understand what enables to learn useful representations, we systematically …

Cross-Modal RetrievalRepresentation LearningRetrieval

Self-Supervised Adversarial Hashing Networks for Cross-Modal Retrieval

2018-04-04 · CVPR 2018 6 · Chao Li, Cheng Deng, Ning li, Wei Liu 외

Thanks to the success of deep learning, cross-modal retrieval has made significant progress recently. However, there still remains a crucial bottleneck: how to bridge the modality gap to further enhance the retrieval acc…

Cross-Modal RetrievalRetrieval

Multimodal Clustering Networks for Self-supervised Learning from Unlabeled Videos

2021-04-26 · ICCV 2021 10 · Brian Chen, Andrew Rouditchenko, Kevin Duarte, Hilde Kuehne 외

Multimodal self-supervised learning is getting more and more attention as it allows not only to train large networks without human supervision but also to search and retrieve data across various modalities. In this conte…

Action LocalizationClusteringContrastive LearningLong Video Retrieval (Background Removed)+5

Self-Supervised Modality-Invariant and Modality-Specific Feature Learning for 3D Objects

2021-09-29 · Longlong Jing, Zhimin Chen, Bing Li, YingLi Tian

While most existing self-supervised 3D feature learning methods mainly focus on point cloud data, this paper explores the inherent multimodal attributes of 3D objects. We propose to jointly learn effective features from …

3D Object RecognitionCross-Modal RetrievalObject RecognitionRetrieval