paper-with-me

홈 › Papers

From Behavioral Performance to Internal Competence: Interpreting Vision-Language Models with VLM-Lens

2025-10-02 · Hala Sheta, Eric Huang, Shuyu Wu, Ilia Alenabi, Jiajun Hong, Ryker Lin, Ruoxi Ning, Daniel Wei, Jialin Yang, Jiawei Zhou, Ziqiao Ma, Freda Shi arxiv

We introduce VLM-Lens, a toolkit designed to enable systematic benchmarking, analysis, and interpretation of vision-language models (VLMs) by supporting the extraction of intermediate outputs from any layer during the forward pass of open-source VLMs. VLM-Lens provides a unified, YAML-configurable interface that abstracts away model-specific complexities and supports user-friendly operation across diverse VLMs. It currently supports 16 state-of-the-art base VLMs and their over 30 variants, and is extensible to accommodate new models without changing the core logic. The toolkit integrates easily with various interpretability and analysis methods. We demonstrate its usage with two simple analytical experiments, revealing systematic differences in the hidden representations of VLMs across layers and target concepts. VLM-Lens is released as an open-sourced project to accelerate community efforts in understanding and improving VLMs.

📄 PDF Abstract BibTeX arXiv:2510.02292

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Pipeline to Assess Merging Methods via Behavior and Internals

2025-09-23 · Yutaro Sigrist, Andreas Waldis arxiv

Merging methods combine the weights of multiple language models (LMs) to leverage their capacities, such as for domain adaptation. While existing studies investigate merged models from a solely behavioral perspective, we…

Domain Adaptation

M3-BENCH: Process-Aware Evaluation of LLM Agents' Social Behaviors in Mixed-Motive Games

2026-01-13 · Sixiong Xie, Zhuofan Shi, Haiyang Shen, Yun Ma 외 arxiv

Existing benchmarks for LLM agents' social behavior typically focus on a single capability dimension and evaluate only behavioral outcomes, overlooking process signals from reasoning and communication. We present M3-BENC…

Evaluating Pragmatic Reasoning in Large Language Models: Evidence from Scalar Diversity

2026-05-09 · Ye-eun Cho arxiv

Evaluating pragmatic reasoning in large language models (LLMs) remains challenging because model behavior can vary depending on evaluation methods. Previous studies suggest that prompt-based judgments may diverge from mo…

What Matters to an LLM? Behavioral and Computational Evidences from Summarization

2026-01-31 · Yongxin Zhou, Changshun Wu, Philippe Mulhem, Didier Schwab 외 arxiv

Large Language Models (LLMs) are now state-of-the-art at summarization, yet the internal notion of importance that drives their information selections remains hidden. We propose to investigate this by combining behaviora…

Estimating Presentation Competence using Multimodal Nonverbal Behavioral Cues

2021-05-06 · Ömer Sümer, Cigdem Beyan, Fabian Ruth, Olaf Kramer 외

Public speaking and presentation competence plays an essential role in many areas of social interaction in our educational, professional, and everyday life. Since our intention during a speech can differ from what is act…