An Evaluation of Image-Based Verb Prediction Models against Human Eye-Tracking Data
Recent research in language and vision has developed models for predicting and disambiguating verbs from images. Here, we ask whether the predictions made by such models correspond to human intuitions about visual verbs. We show that the image regions a verb prediction model identifies as salient for a given verb correlate with the regions fixated by human observers performing a verb classification task.
Code (0)
등록된 구현이 없습니다.
Tasks
General ClassificationQuestion AnsweringVisual Question Answering (VQA)Word Sense DisambiguationSimilar Papers 제목 키워드 기반
VideoNorms: Benchmarking Cultural Awareness of Video Language Models
As Video Large Language Models (VideoLLMs) are deployed globally, it is important to assess their ability to reason across cultural contexts. To advance cultural norm awareness evaluation in VideoLLMs, we introduce Video…
Saliency Map Verbalization: Comparing Feature Importance Representations from Model-free and Instruction-based Methods
Saliency maps can explain a neural model's predictions by identifying important input features. They are difficult to interpret for laypeople, especially for instances with many features. In order to make them more acces…
Abstractive Text SummarizationFeature ImportanceSentiment Analysistext-classification+2A Psycholinguistic Evaluation of Language Models' Sensitivity to Argument Roles
We present a systematic evaluation of large language models' sensitivity to argument roles, i.e., who did what to whom, by replicating psycholinguistic studies on human argument role processing. In three experiments, we …
SensitivitySentenceComparing Apples to Oranges: LLM-powered Multimodal Intention Prediction in an Object Categorization Task
Human intention-based systems enable robots to perceive and interpret user actions to interact with humans and adapt to their behavior proactively. Therefore, intention prediction is pivotal in creating a natural interac…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Object Categorizationspeech-recognition+2Towards Robust Evaluation of Visual Activity Recognition: Resolving Verb Ambiguity with Sense Clustering
Evaluating visual activity recognition systems is challenging due to inherent ambiguities in verb semantics and image interpretation. When describing actions in images, synonymous verbs can refer to the same event (e.g.,…
Activity Recognition