paper-with-me

홈 › Papers

InterVLS: Interactive Model Understanding and Improvement with Vision-Language Surrogates

2023-11-06 · Jinbin Huang, Wenbin He, Liang Gou, Liu Ren, Chris Bryan

Deep learning models are widely used in critical applications, highlighting the need for pre-deployment model understanding and improvement. Visual concept-based methods, while increasingly used for this purpose, face challenges: (1) most concepts lack interpretability, (2) existing methods require model knowledge, often unavailable at run time. Additionally, (3) there lacks a no-code method for post-understanding model improvement. Addressing these, we present InterVLS. The system facilitates model understanding by discovering text-aligned concepts, measuring their influence with model-agnostic linear surrogates. Employing visual analytics, InterVLS offers concept-based explanations and performance insights. It enables users to adjust concept influences to update a model, facilitating no-code model improvement. We evaluate InterVLS in a user study, illustrating its functionality with two scenarios. Results indicates that InterVLS is effective to help users identify influential concepts to a model, gain insights and adjust concept influence to improve the model. We conclude with a discussion based on our study results.

📄 PDF Abstract BibTeX arXiv:2311.03547

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving

2025-05-13 · Zongchuang Zhao, Haoyu Fu, Dingkang Liang, Xin Zhou 외

The Large Visual-Language Models (LVLMs) have significantly advanced image understanding. Their comprehension and reasoning capabilities enable promising applications in autonomous driving scenarios. However, existing re…

3D visual groundingAutonomous DrivingLanguage ModelingLanguage Modelling+3

SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL

2025-12-03 · Siyi Chen, Mikaela Angelina Uy, Chan Hee Song, Faisal Ladhak 외 arxiv

Vision Language Models (VLMs) demonstrate strong qualitative visual understanding, but struggle with metrically precise spatial reasoning required for embodied applications. The agentic paradigm promises that VLMs can us…

Reinforcement LearningSpatial Reasoning

Toward Interactive Regional Understanding in Vision-Large Language Models

2024-03-27 · Jungbeom Lee, Sanghyuk Chun, Sangdoo Yun

Recent Vision-Language Pre-training (VLP) models have demonstrated significant advancements. Nevertheless, these models heavily rely on image-text pairs that capture only coarse and global information of an image, leadin…

Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

2024-03-19 · Zuyan Liu, Yuhao Dong, Yongming Rao, Jie zhou 외

In the realm of vision-language understanding, the proficiency of models in interpreting and reasoning over visual content has become a cornerstone for numerous applications. However, it is challenging for the visual enc…

Instruction Followingvisual instruction followingVisual Question Answering

Scenes as Tokens: Multi-Scale Normal Distributions Transform Tokenizer for General 3D Vision-Language Understanding

2025-11-26 · Yutao Tang, Cheng Zhao, Gaurav Mittal, Rohith Kukkala 외 arxiv

Recent advances in 3D vision-language models (VLMs) highlight a strong potential for 3D scene understanding and reasoning. However, effectively tokenizing 3D scenes into holistic scene tokens, and leveraging these tokens…

Visual Question Answering3D dense captioningScene UnderstandingPoint Clouds