paper-with-me

Papers

VLMs May Not Globally Enhance Human Alignment over LLMs During Natural Reading

2026-05-27 · Jinzhou Wu, Zhengwu Ma, Jixing Li, Baoping Tang, Zitong Lu arxiv

Large language models (LLMs) have become increasingly useful computational models of human language processing, but it remains unclear whether vision-language learning makes text representations more human-like during natural reading. Here, we address this question by comparing tightly matched LLM and vision-language model (VLM) pairs under a strictly text-only setting, allowing us to isolate the effect of multimodal training history from online visual input or cross-modal fusion. We evaluate model alignment with a human natural-reading dataset that includes whole-cortex fMRI responses and synchronized eye-tracking saccades. Our findings demonstrate that multimodal pretraining may not confer a uniform, global advantage in human alignment during natural reading, indicating that language-internal representations remain the key factor for modeling human text processing. However, the VLM advantage could emerge more selectively when sentences contain stronger visual semantic content, with converging evidence from both fMRI and eye-movement alignments. Together, our findings provide a controlled in silico framework for testing how visual learning history shapes model-human alignment of language processing, suggesting that multimodal pretraining contributes selectively rather than globally to human-like language representations during natural reading.

📄 PDF Abstract BibTeX arXiv:2605.28818

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Empowering Semantic-Sensitive Underwater Image Enhancement with VLM

2026-03-13 · Guodong Fan, Shengning Zhou, Genji Yuan, Huiyu Li 외 arxiv

In recent years, learning-based underwater image enhancement (UIE) techniques have rapidly evolved. However, distribution shifts between high-quality enhanced outputs and natural images can hinder semantic cue extraction…

Image ReconstructionImage Enhancement

Improving Alignment in LVLMs with Debiased Self-Judgment

2025-08-28 · Sihan Yang, Chenhang Cui, Zihao Zhao, Yiyang Zhou 외 arxiv

The rapid advancements in Large Language Models (LLMs) and Large Visual-Language Models (LVLMs) have opened up new opportunities for integrating visual and linguistic modalities. However, effectively aligning these modal…

DRESS: Instructing Large Vision-Language Models to Align and Interact with Humans via Natural Language Feedback

2023-11-16 · CVPR 2024 1 · Yangyi Chen, Karan Sikka, Michael Cogswell, Heng Ji 외

We present DRESS, a large vision language model (LVLM) that innovatively exploits Natural Language feedback (NLF) from Large Language Models to enhance its alignment and interactions by addressing two key limitations in …

Language Modelling

ChartGaze: Enhancing Chart Understanding in LVLMs with Eye-Tracking Guided Attention Refinement

2025-09-16 · Ali Salamatian, Amirhossein Abaskohi, Wan-Cyuan Fan, Mir Rayat Imtiaz Hossain 외 arxiv

Charts are a crucial visual medium for communicating and representing information. While Large Vision-Language Models (LVLMs) have made progress on chart question answering (CQA), the task remains challenging, particular…

Chart Question Answering

MMedPO: Aligning Medical Vision-Language Models with Clinical-Aware Multimodal Preference Optimization

2024-12-09 · Kangyu Zhu, Peng Xia, Yun Li, Hongtu Zhu 외

The advancement of Large Vision-Language Models (LVLMs) has propelled their application in the medical field. However, Medical LVLMs (Med-LVLMs) encounter factuality challenges due to modality misalignment, where the mod…

Visual Question Answering (VQA)