paper-with-me

홈 › Papers

WaveMind: Towards a Conversational EEG Foundation Model Aligned to Textual and Visual Modalities

2025-09-26 · Ziyi Zeng, Zhenyang Cai, Yixi Cai, Xidong Wang, Junying Chen, Rongsheng Wang, Yipeng Liu, Siqi Cai, Benyou Wang, Zhiguo Zhang, Haizhou Li arxiv

Electroencephalography (EEG) interpretation using multimodal large language models (MLLMs) offers a novel approach for analyzing brain signals. However, the complex nature of brain activity introduces critical challenges: EEG signals simultaneously encode both cognitive processes and intrinsic neural states, creating a mismatch in EEG paired-data modality that hinders effective cross-modal representation learning. Through a pivot investigation, we uncover complementary relationships between these modalities. Leveraging this insight, we propose mapping EEG signals and their corresponding modalities into a unified semantic space to achieve generalized interpretation. To fully enable conversational capabilities, we further introduce WaveMind-Instruct-338k, the first cross-task EEG dataset for instruction tuning. The resulting model demonstrates robust classification accuracy while supporting flexible, open-ended conversations across four downstream tasks, thereby offering valuable insights for both neuroscience research and the development of general-purpose EEG models.

📄 PDF Abstract BibTeX arXiv:2510.00032

Code (0)

등록된 구현이 없습니다.

Tasks

Representation Learning

Similar Papers 제목 키워드 기반

Beyond Words: Multimodal LLM Knows When to Speak

2025-05-20 · Zikai Liao, Yi Ouyang, Yi-Lun Lee, Chen-Ping Yu 외

While large language model (LLM)-based chatbots have demonstrated strong capabilities in generating coherent and contextually relevant responses, they often struggle with understanding when to speak, particularly in deli…

Large Language Model

Show and Guide: Instructional-Plan Grounded Vision and Language Model

2024-09-27 · Diogo Glória-Silva, David Semedo, João Magalhães

Guiding users through complex procedural plans is an inherently multimodal task in which having visually illustrated plan steps is crucial to deliver an effective plan guidance. However, existing works on plan-following …

Language ModelingLanguage ModellingMoment Retrieval

Insect-Foundation: A Foundation Model and Large Multimodal Dataset for Vision-Language Insect Understanding

2025-02-14 · Thanh-Dat Truong, Hoang-Quan Nguyen, Xuan-Bac Nguyen, Ashley Dowling 외

Multimodal conversational generative AI has shown impressive capabilities in various vision and language understanding through learning massive text-image data. However, current conversational models still lack knowledge…

General KnowledgeQuestion AnsweringSelf-Supervised Learning

Augmenting a Large Language Model with a Combination of Text and Visual Data for Conversational Visualization of Global Geospatial Data

2025-01-16 · Omar Mena, Alexandre Kouyoumdjian, Lonni Besançon, Michael Gleicher 외

We present a method for augmenting a Large Language Model (LLM) with a combination of text and visual data to enable accurate question answering in visualization of scientific data, making conversational visualization po…

Data InteractionDescriptiveLanguage ModelingLanguage Modelling+2

Biomedical Visual Instruction Tuning with Clinician Preference Alignment

2024-06-19 · Hejie Cui, Lingjun Mao, Xin Liang, Jieyu Zhang 외

Recent advancements in multimodal foundation models have showcased impressive capabilities in understanding and reasoning with visual and textual information. Adapting these foundation models trained for general usage to…

Instruction FollowingVisual Question Answering (VQA)