Self-supervised speech representation and contextual text embedding for match-mismatch classification with EEG recording
Relating speech to EEG holds considerable importance but is challenging. In this study, a deep convolutional network was employed to extract spatiotemporal features from EEG data. Self-supervised speech representation and contextual text embedding were used as speech features. Contrastive learning was used to relate EEG features to speech features. The experimental results demonstrate the benefits of using self-supervised speech representation and contextual text embedding. Through feature fusion and model ensemble, an accuracy of 60.29% was achieved, and the performance was ranked as No.2 in Task 1 of the Auditory EEG Challenge (ICASSP 2024). The code to implement our work is available on Github: https://github.com/bobwangPKU/EEG-Stimulus-Match-Mismatch.
Code (1)
Tasks
Contrastive LearningEEGMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
DM-Codec: Distilling Multimodal Representations for Speech Tokenization
Recent advancements in speech-language models have yielded significant improvements in speech tokenization and synthesis. However, effectively mapping the complex, multidimensional attributes of speech into discrete toke…
Self-Supervised LearningSpeech TokenizationAn empirical study on speech restoration guided by self supervised speech representation
Enhancing speech quality is an indispensable yet difficult task as it is often complicated by a range of degradation factors. In addition to additive noise, reverberation, clipping, and speech attenuation can all adverse…
Representation LearningSpeech Representation LearningEfficient Self-supervised Learning with Contextualized Target Representations for Vision, Speech and Language
Current self-supervised learning algorithms are often modality-specific and require large amounts of computational resources. To address these issues, we increase the training efficiency of data2vec, a learning objective…
Decoderimage-classificationImage ClassificationNatural Language Understanding+3FuseCodec: Semantic-Contextual Fusion and Supervision for Neural Codecs
Speech tokenization enables discrete representation and facilitates speech language modeling. However, existing neural codecs capture low-level acoustic features, overlooking the semantic and contextual cues inherent to …
Representation LearningSpeech SynthesisAV-data2vec: Self-supervised Learning of Audio-Visual Speech Representations with Contextualized Target Representations
Self-supervision has shown great potential for audio-visual speech recognition by vastly reducing the amount of labeled data required to build good systems. However, existing methods are either not entirely end-to-end or…
Audio-Visual Speech RecognitionSelf-Supervised Learningspeech-recognitionSpeech Recognition+1