paper-with-me

홈 › Papers

Visualizing the Invisible: Generative Visual Grounding Empowers Universal EEG Understanding in MLLMs

2026-05-18 · Jun-Yu Pan, Yansen Wang, Enze Zhang, Bao-Liang Lu, Wei-Long Zheng, Dongsheng Li arxiv

Leveraging the universal representations of pre-trained LLMs and MLLMs offers a promising path toward brain foundation models. However, visually-evoked EEG datasets remain scarce, leading existing methods to align neural signals mainly with abstract text, a lossy translation that may discard fine-grained perceptual information encoded in brain activity. We propose Generative Visual Grounding (GVG), a framework that visualizes the invisible by using an EEG-to-image generative model as a visual translator. Instead of forcing EEG into text alone, GVG hallucinates instance-specific proxy images for non-visual EEG, providing structured visual contexts that allow MLLMs to exploit their visual priors for clinical-state interpretation. We validate this idea on two MLLM backbones, GVG-X-Omni and GVG-Janus. Image-only alignment is already competitive: the lightweight GVG-X-Omni matches 1.7B-parameter text-aligned baselines while tuning only 170M parameters on a frozen 7B backbone. We further extend GVG-Janus with trimodal Image+Text alignment, where text supplies categorical semantic anchors and visual proxies enrich neural representations with perceptual details. Experiments show consistent gains in EEG understanding and visual generation, suggesting visual proxy grounding as an effective complement to textual alignment.

📄 PDF Abstract BibTeX arXiv:2605.18172

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Grounding

Similar Papers 제목 키워드 기반

Visualizing the Invisible: Occluded Vehicle Segmentation and Recovery

2019-07-22 · ICCV 2019 10 · Xiaosheng Yan, Yuanlong Yu, Feigege Wang, Wenxi Liu 외

In this paper, we propose a novel iterative multi-task framework to complete the segmentation mask of an occluded vehicle and recover the appearance of its invisible parts. In particular, to improve the quality of the se…

Segmentation

VisualTrap: A Stealthy Backdoor Attack on GUI Agents via Visual Grounding Manipulation

2025-07-09 · Ziang Ye, Yang Zhang, Wentao Shi, Xiaoyu You 외

Graphical User Interface (GUI) agents powered by Large Vision-Language Models (LVLMs) have emerged as a revolutionary approach to automating human-machine interactions, capable of autonomously operating personal devices …

Backdoor AttackVisual Grounding

VGGDrive: Empowering Vision-Language Models with Cross-View Geometric Grounding for Autonomous Driving

2026-02-24 · Jie Wang, Guang Li, Zhijian Huang, Chenxu Dang 외 arxiv

The significance of cross-view 3D geometric modeling capabilities for autonomous driving is self-evident, yet existing Vision-Language Models (VLMs) inherently lack this capability, resulting in their mediocre performanc…

Trajectory PlanningAutonomous Driving

Invisible Strings: Revealing Latent Dancer-to-Dancer Interactions with Graph Neural Networks

2025-03-04 · Luis Vitor Zerkowski, Zixuan Wang, Ilya Vidrin, Mariel Pettee

Dancing in a duet often requires a heightened attunement to one's partner: their orientation in space, their momentum, and the forces they exert on you. Dance artists who work in partnered settings might have a strong em…

Mapping Patient Trajectories: Understanding and Visualizing Sepsis Prognostic Pathways from Patients Clinical Narratives

2024-07-20 · Sudeshna Jana, Tirthankar Dasgupta, Lipika Dey

In recent years, healthcare professionals are increasingly emphasizing on personalized and evidence-based patient care through the exploration of prognostic pathways. To study this, structured clinical variables from Ele…