paper-with-me

Papers

Independent Feature Enhanced Crossmodal Fusion for Match-Mismatch Classification of Speech Stimulus and EEG Response

2024-10-19 · Shitong Fan, Wenbo Wang, Feiyang Xiao, Shiheng Zhang, Qiaoxi Zhu, Jian Guan

It is crucial for auditory attention decoding to classify matched and mismatched speech stimuli with corresponding EEG responses by exploring their relationship. However, existing methods often adopt two independent networks to encode speech stimulus and EEG response, which neglect the relationship between these signals from the two modalities. In this paper, we propose an independent feature enhanced crossmodal fusion model (IFE-CF) for match-mismatch classification, which leverages the fusion feature of the speech stimulus and the EEG response to achieve auditory EEG decoding. Specifically, our IFE-CF contains a crossmodal encoder to encode the speech stimulus and the EEG response with a two-branch structure connected via crossmodal attention mechanism in the encoding process, a multi-channel fusion module to fuse features of two modalities by aggregating the interaction feature obtained from the crossmodal encoder and the independent feature obtained from the speech stimulus and EEG response, and a predictor to give the matching result. In addition, the causal mask is introduced to consider the time delay of the speech-EEG pair in the crossmodal encoder, which further enhances the feature representation for match-mismatch classification. Experiments demonstrate our method's effectiveness with better classification accuracy, as compared with the baseline of the Auditory EEG Decoding Challenge 2023.

📄 PDF Abstract BibTeX arXiv:2410.15078

Code (0)

등록된 구현이 없습니다.

Tasks

EEGEeg Decoding

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

Attention Is Not Enough: Mitigating the Distribution Discrepancy in Asynchronous Multimodal Sequence Fusion

2021-01-01 · ICCV 2021 10 · Tao Liang, Guosheng Lin, Lei Feng, Yan Zhang 외

Videos flow as the mixture of language, acoustic, and vision modalities. A thorough video understanding needs to fuse time-series data of different modalities for prediction. Due to the variable receiving frequency f…

Time SeriesTime Series AnalysisVideo Understanding

WCCNet: Wavelet-integrated CNN with Crossmodal Rearranging Fusion for Fast Multispectral Pedestrian Detection

2023-08-02 · Xingjian Wang, Li Chai, Jiming Chen, Zhiguo Shi

Multispectral pedestrian detection achieves better visibility in challenging conditions and thus has a broad application in various tasks, for which both the accuracy and computational cost are of paramount importance. M…

Computational EfficiencyPedestrian Detection

Cross Spatial Temporal Fusion Attention for Remote Sensing Object Detection via Image Feature Matching

2025-07-25 · Abu Sadat Mohammad Salehin Amit, Xiaoli Zhang, Md Masum Billa Shagar, Zhaojun Liu 외 arxiv

Effectively describing features for cross-modal remote sensing image matching remains a challenging task due to the significant geometric and radiometric differences between multimodal images. Existing methods primarily …

Computational EfficiencyObject DetectionImage Matching

Unimodal and Crossmodal Refinement Network for Multimodal Sequence Fusion

2021-11-01 · EMNLP 2021 11 · Xiaobao Guo, Adams Kong, Huan Zhou, Xianfeng Wang 외

Effective unimodal representation and complementary crossmodal representation fusion are both important in multimodal representation learning. Prior works often modulate one modal feature to another straightforwardly and…

Representation Learning

CrossModalityDiffusion: Multi-Modal Novel View Synthesis with Unified Intermediate Representation

2025-01-16 · Alex Berian, Daniel Brignac, JhihYang Wu, Natnael Daba 외

Geospatial imaging leverages data from diverse sensing modalities-such as EO, SAR, and LiDAR, ranging from ground-level drones to satellite views. These heterogeneous inputs offer significant opportunities for scene unde…

Novel View SynthesisScene Understanding