paper-with-me

Papers

Responses Fall Short of Understanding: Revealing the Gap between Internal Representations and Responses in Visual Document Understanding

2026-04-06 · Haruka Kawasaki, Ryota Tanaka, Kyosuke Nishida arxiv

Visual document understanding (VDU) is a challenging task for large vision language models (LVLMs), requiring the integration of visual perception, text recognition, and reasoning over structured layouts. Although recent LVLMs have shown progress on VDU benchmarks, their performance is typically evaluated based on generated responses, which may not necessarily reflect whether the model has actually captured the required information internally. In this paper, we investigate how information required to solve VDU tasks is represented across different layers of LLMs within LVLMs using linear probing. Our study reveals that (1) there is a clear gap between internal representations and generated responses, and (2) information required to solve the task is often encoded more linearly from intermediate layers than from the final layer. Motivated by these findings, we explore fine-tuning strategies that target intermediate layers. Experiments show that fine-tuning intermediate layers improves both linear probing accuracy and response accuracy while narrowing the gap.

📄 PDF Abstract BibTeX arXiv:2604.04411

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Design Challenges for a Multi-Perspective Search Engine

2021-12-15 · Findings (NAACL) 2022 7 · Sihao Chen, Siyi Liu, Xander Uyttendaele, Yi Zhang 외

Many users turn to document retrieval systems (e.g. search engines) to seek answers to controversial questions. Answering such user queries usually require identifying responses within web documents, and aggregating the …

Natural Language UnderstandingRetrieval

A Residual Bootstrap for Conditional Expected Shortfall

2018-11-26

This paper studies a fixed-design residual bootstrap method for the two-step estimator of Francq and Zako\"ian (2015) associated with the conditional Expected Shortfall. For a general class of volatility models the boots…

valid

ClustOpt: A Clustering-based Approach for Representing and Visualizing the Search Dynamics of Numerical Metaheuristic Optimization Algorithms

2025-07-03 · Gjorgjina Cenikj, Gašper Petelin, Tome Eftimov arxiv

Understanding the behavior of numerical metaheuristic optimization algorithms is critical for advancing their development and application. Traditional visualization techniques, such as convergence plots, trajectory mappi…

Global analysis reveals persistent shortfalls and regional differences in availability of foods needed for health

2024-01-02 · Leah Costlow, Anna Herforth, Timothy B. Sulser, Nicola Cenacchi 외

Most people around the world still lack access to sufficient quantities of all food groups needed for an active and healthy life. This study traces historical and projected changes in global food systems toward alignment…

MultiVox: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions

2025-07-14 · Ramaneswaran Selvakumar, Ashish Seth, Nishit Anand, Utkarsh Tyagi 외 arxiv

The rapid progress of Large Language Models (LLMs) has empowered omni models to act as voice assistants capable of understanding spoken dialogues. These models can process multimodal inputs beyond text, such as speech an…