paper-with-me

Papers

LoGra-Med: Long Context Multi-Graph Alignment for Medical Vision-Language Model

2024-10-03 · Duy M. H. Nguyen, Nghiem T. Diep, Trung Q. Nguyen, Hoang-Bao Le, Tai Nguyen, Tien Nguyen, TrungTin Nguyen, Nhat Ho, Pengtao Xie, Roger Wattenhofer, James Zhou, Daniel Sonntag, Mathias Niepert

State-of-the-art medical multi-modal large language models (med-MLLM), like LLaVA-Med or BioMedGPT, leverage instruction-following data in pre-training. However, those models primarily focus on scaling the model size and data volume to boost performance while mainly relying on the autoregressive learning objectives. Surprisingly, we reveal that such learning schemes might result in a weak alignment between vision and language modalities, making these models highly reliant on extensive pre-training datasets - a significant challenge in medical domains due to the expensive and time-consuming nature of curating high-quality instruction-following instances. We address this with LoGra-Med, a new multi-graph alignment algorithm that enforces triplet correlations across image modalities, conversation-based descriptions, and extended captions. This helps the model capture contextual meaning, handle linguistic variability, and build cross-modal associations between visuals and text. To scale our approach, we designed an efficient end-to-end learning scheme using black-box gradient estimation, enabling faster LLaMa 7B training. Our results show LoGra-Med matches LLAVA-Med performance on 600K image-text pairs for Medical VQA and significantly outperforms it when trained on 10% of the data. For example, on VQA-RAD, we exceed LLAVA-Med by 20.13% and nearly match the 100% pre-training score (72.52% vs. 72.64%). We also surpass SOTA methods like BiomedGPT on visual chatbots and RadFM on zero-shot image classification with VQA, highlighting the effectiveness of multi-graph alignment.

📄 PDF Abstract BibTeX arXiv:2410.02615

Code (0)

등록된 구현이 없습니다.

Tasks

image-classificationImage ClassificationInstruction FollowingLanguage ModelingLanguage ModellingTripletVisual Question Answering (VQA)Zero-Shot Image Classification

Methods 이 논문이 사용한 방법론

LLaMA LLaMA is a collection of foundation language models ranging from 7B to 65B parameters. It is based on the transformer architecture with various improvements that were…
Focus 설명 없음

Similar Papers 제목 키워드 기반

MAD: Multi-Alignment MEG-to-Text Decoding

2024-06-03 · Yiqian Yang, Hyejeong Jo, Yiqun Duan, Qiang Zhang 외

Deciphering language from brain activity is a crucial task in brain-computer interface (BCI) research. Non-invasive cerebral signaling techniques including electroencephalography (EEG) and magnetoencephalography (MEG) ar…

Brain Computer InterfaceEEGText Generation

Tac-DINO: Learning Vision-Tactile Features with Patch Alignment

2026-06-10 · Hong Li, Yankang Dong, Yue Xu, Yihan Tang 외 arxiv

Touch is the primary medium through which humans interact with the environment. Currently, tactile learning mainly focuses on image-level pretraining or alignment. However, tactile signals correspond to local object cont…

Representation Learning

LION: A Clifford Neural Paradigm for Multimodal-Attributed Graph Learning

2026-01-29 · Xunkai Li, Zhengyu Wu, Zekai Chen, Henan Sun 외 arxiv

Recently, the rapid advancement of multimodal domains has driven a data-centric paradigm shift in graph ML, transitioning from text-attributed to multimodal-attributed graphs. This advancement significantly enhances data…

Graph Learning

OPENXRD: A Comprehensive Benchmark Framework for LLM/MLLM XRD Question Answering

2025-07-12 · Ali Vosoughi, Ayoub Shahnazari, Yufeng Xi, Zeliang Zhang 외 arxiv

We introduce OPENXRD, a comprehensive benchmarking framework for evaluating large language models (LLMs) and multimodal LLMs (MLLMs) in crystallography question answering. The framework measures context assimilation, or …

Question Answering

Contextual Minimum-Norm Estimates (CMNE): A Deep Learning Method for Source Estimation in Neuronal Networks

2019-09-05 · Christoph Dinh, John GW Samuelsson, Alexander Hunold, Matti S Hämäläinen 외

Magnetoencephalography (MEG) and Electroencephalography (EEG) source estimates have thus far mostly been derived sample by sample, i.e., independent of each other in time. However, neuronal assemblies are heavily interco…

EEGElectroencephalogram (EEG)