paper-with-me

Papers

EgoArgus: Benchmarking VLMs as Situational Assistants for Modality-Grounded User Supports

2026-08-26 · Yu-Chien Tang, Yu-Hsiang Liu, An-Zi Yen arxiv

VLMs are increasingly positioned as daily assistants that perceive first-person environments, follow user dialogue, and decide how to help. Existing egocentric benchmarks mainly evaluate visual understanding in isolation, leaving open whether models can arbitrate between visual evidence and user-provided language when the two are helpful, irrelevant, or conflicting. We introduce EgoArgus, a human-annotated dataset for evaluating egocentric assistants on understanding and decision tasks in five dialogue-video daily scenarios. Our results demonstrate that it is still challenging for current VLMs as reliable egocentric assistants, which requires identifying which modality is trustworthy and deciding when intervention is warranted. Deeper analysis also shows that existing modality bias mitigation methods are quite restricted to enhance performance, providing insights to aid practioners into the deployment of current VLMs as daily assistants.

📄 PDF Abstract BibTeX arXiv:2608.25561

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Vision-Language Models under Cultural and Inclusive Considerations

2024-07-08 · Antonia Karamolegkou, Phillip Rust, Yong Cao, Ruixiang Cui 외

Large vision-language models (VLMs) can assist visually impaired people by describing images from their daily lives. Current evaluation datasets may not reflect diverse cultural user backgrounds or the situational contex…

HallucinationSurvey

Continual Vision-Language Learning for Remote Sensing: Benchmarking and Analysis

2026-04-01 · Xingxing Weng, Ruifeng Ni, Chao Pang, XiangYu Hao 외 arxiv

Current remote sensing vision-language models (RS VLMs) demonstrate impressive performance in image interpretation but rely on static training data, limiting their ability to accommodate continuously emerging sensing mod…

Continual Learning

AutoBench-V: Can Large Vision-Language Models Benchmark Themselves?

2024-10-28 · Han Bao, Yue Huang, Yanbo Wang, Jiayi Ye 외

Large Vision-Language Models (LVLMs) have become essential for advancing the integration of visual and linguistic information. However, the evaluation of LVLMs presents significant challenges as the evaluation benchmark …

BenchmarkingQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs

2024-07-05 · Rudolf Laine, Bilal Chughtai, Jan Betley, Kaivalya Hariharan 외

AI assistants such as ChatGPT are trained to respond to users by saying, "I am a large language model". This raises questions. Do such models know that they are LLMs and reliably act on this knowledge? Are they aware of …

General KnowledgeInstruction FollowingLanguage ModellingLarge Language Model+2

Beyond Symmetric Alignment: Spectral Diagnostics of Modality Imbalance in Vision-Language Models in the Medical Domain

2026-06-03 · Alessandro Gambetti, Qiwei Han, Cláudia Soares, Hong Shen arxiv

Vision-Language Models (VLMs) struggle when applied to medical image-text data, yet the tools available to diagnose this failure remain limited. Existing representation alignment metrics are symmetric, collapsing both mo…