paper-with-me

홈 › Papers

To Trust Or Not To Trust Your Vision-Language Model's Prediction

2025-05-29 · Hao Dong, Moru Liu, Jian Liang, Eleni Chatzi, Olga Fink

Vision-Language Models (VLMs) have demonstrated strong capabilities in aligning visual and textual modalities, enabling a wide range of applications in multimodal understanding and generation. While they excel in zero-shot and transfer learning scenarios, VLMs remain susceptible to misclassification, often yielding confident yet incorrect predictions. This limitation poses a significant risk in safety-critical domains, where erroneous predictions can lead to severe consequences. In this work, we introduce TrustVLM, a training-free framework designed to address the critical challenge of estimating when VLM's predictions can be trusted. Motivated by the observed modality gap in VLMs and the insight that certain concepts are more distinctly represented in the image embedding space, we propose a novel confidence-scoring function that leverages this space to improve misclassification detection. We rigorously evaluate our approach across 17 diverse datasets, employing 4 architectures and 2 VLMs, and demonstrate state-of-the-art performance, with improvements of up to 51.87% in AURC, 9.14% in AUROC, and 32.42% in FPR95 compared to existing baselines. By improving the reliability of the model without requiring retraining, TrustVLM paves the way for safer deployment of VLMs in real-world applications. The code will be available at https://github.com/EPFL-IMOS/TrustVLM.

📄 PDF Abstract BibTeX arXiv:2505.23745

Code (1)

epfl-imos/trustvlm 공식 구현 pytorch

Tasks

Transfer Learning

Similar Papers 제목 키워드 기반

Know Your Model (KYM): Increasing Trust in AI and Machine Learning

2021-05-31 · Mary Roszel, Robert Norvill, Jean Hilger, Radu State

The widespread utilization of AI systems has drawn attention to the potential impacts of such systems on society. Of particular concern are the consequences that prediction errors may have on real-world scenarios, and th…

BIG-bench Machine Learning

Can I Trust Your Answer? Visually Grounded Video Question Answering

2023-09-04 · CVPR 2024 1 · Junbin Xiao, Angela Yao, Yicong Li, Tat Seng Chua

We study visually grounded VideoQA in response to the emerging trends of utilizing pretraining techniques for video-language understanding. Specifically, by forcing vision-language models (VLMs) to answer questions and s…

Grounded Video Question AnsweringQuestion AnsweringVideo GroundingVideo Question Answering+1

Fool Your (Vision and) Language Model With Embarrassingly Simple Permutations

2023-10-02 · Yongshuo Zong, Tingyang Yu, Ruchika Chavhan, Bingchen Zhao 외

Large language and vision-language models are rapidly being deployed in practice thanks to their impressive capabilities in instruction following, in-context learning, and so on. This raises an urgent need to carefully a…

In-Context LearningInstruction FollowingLanguage ModelingLanguage Modelling+3

Trust the Model Where It Trusts Itself -- Model-Based Actor-Critic with Uncertainty-Aware Rollout Adaption

2024-05-29 · Bernd Frauenknecht, Artur Eisele, Devdutt Subhasish, Friedrich Solowjow 외

Dyna-style model-based reinforcement learning (MBRL) combines model-free agents with predictive transition models through model-based rollouts. This combination raises a critical question: 'When to trust your model?'; i.…

modelModel-based Reinforcement LearningMuJoCo

Trusting Your AI Agent Emotionally and Cognitively: Development and Validation of a Semantic Differential Scale for AI Trust

2024-07-25 · Ruoxi Shang, Gary Hsieh, Chirag Shah

Trust is not just a cognitive issue but also an emotional one, yet the research in human-AI interactions has primarily focused on the cognitive route of trust development. Recent work has highlighted the importance of st…

AI Agent