paper-with-me

홈 › Papers

TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic Restraint

2026-08-13 · Fnu Pramono, John Cai, Sourabh Kulkarni arxiv

When visual evidence is occluded or chaotic, models should abstain. In this paper, we show that Vision-Language Models (VLMs) can internally distinguish when abstention is required, but fail to express it anyway. We introduce TRAPSBench, a procedurally generated video benchmark of 1,404 matched physics pairs in which a single targeted change renders the outcome undeterminable from the visual evidence. Furthermore, we introduce Penalized Epistemic Calibration Score (PECS), a new robust metric that requires models to both answer correctly when the outcome is knowable, and abstain when the outcome is not. Across 16 VLMs spanning five families, spontaneous restraint is poor: the best PECS is 0.292. The bottleneck is expression, not perception: linear probes decode answerability from hidden states at up to 0.91 AUROC across physics domains; steering a single-layer void direction causally induces or suppresses abstention. Our results replicate across three open-weight families (Qwen, Gemma, LLaVA). The failure is also more pronounced in visual than textual uncertainty: models detect textual impossibility about 4x more readily than missing visual evidence. Closing this representation--output gap likely requires output-stage interventions.

📄 PDF Abstract BibTeX arXiv:2608.13167

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LaCoVL-FER: Landmark-Guided Contrastive Learning Network with Vision-Language Enhancement for Facial Expression Recognition

2026-05-19 · Jiaxin Wang, Muwei Jian, Hui Yu, Junyu Dong 외 arxiv

Facial Expression Recognition (FER) in the wild requires models to identify subtle expression cues under large variations in pose, occlusion, illumination, and identity. Recent FER methods improve robustness by introduci…

Facial Expression RecognitionContrastive Learning

Enhancing Zero-Shot Facial Expression Recognition by LLM Knowledge Transfer

2024-05-29 · Zengqun Zhao, Yu Cao, Shaogang Gong, Ioannis Patras

Current facial expression recognition (FER) models are often designed in a supervised learning manner and thus are constrained by the lack of large-scale facial expression images with high-quality annotations. Consequent…

Facial Expression RecognitionFacial Expression Recognition (FER)Transfer LearningZero-Shot Facial Expression Recognition

SpatialBoost: Enhancing Visual Representation through Language-Guided Reasoning

2026-03-23 · Byungwoo Jeon, Dongyoung Kim, Huiwon Jang, Insoo Kim 외 arxiv

Despite the remarkable success of large-scale pre-trained image representation models (i.e., vision encoders) across various vision tasks, they are predominantly trained on 2D image data and therefore often fail to captu…

Location-Aware Pretraining for Medical Difference Visual Question Answering

2026-03-05 · Denis Musinguzi, Caren Han, Prasenjit Mitra arxiv

Differential medical VQA models compare multiple images to identify clinically meaningful changes and rely on vision encoders to capture fine-grained visual differences that reflect radiologists' comparative diagnostic w…

Visual Question Answering

Cross-aware Early Fusion with Stage-divided Vision and Language Transformer Encoders for Referring Image Segmentation

2024-08-14 · Yubin Cho, Hyunwoo Yu, Suk-Ju Kang

Referring segmentation aims to segment a target object related to a natural language expression. Key challenges of this task are understanding the meaning of complex and ambiguous language expressions and determining the…

cross-modal alignmentImage SegmentationSemantic Segmentation