paper-with-me

홈 › Papers

An Evaluation of Image-Based Verb Prediction Models against Human Eye-Tracking Data

2018-06-01 · NAACL 2018 6 · Sp Gella, ana, Frank Keller

Recent research in language and vision has developed models for predicting and disambiguating verbs from images. Here, we ask whether the predictions made by such models correspond to human intuitions about visual verbs. We show that the image regions a verb prediction model identifies as salient for a given verb correlate with the regions fixated by human observers performing a verb classification task.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

General ClassificationQuestion AnsweringVisual Question Answering (VQA)Word Sense Disambiguation

Similar Papers 제목 키워드 기반

VideoNorms: Benchmarking Cultural Awareness of Video Language Models

2025-10-09 · Nikhil Reddy Varimalla, Yunfei Xu, Arkadiy Saakyan, Meng Fan Wang 외 arxiv

As Video Large Language Models (VideoLLMs) are deployed globally, it is important to assess their ability to reason across cultural contexts. To advance cultural norm awareness evaluation in VideoLLMs, we introduce Video…

Saliency Map Verbalization: Comparing Feature Importance Representations from Model-free and Instruction-based Methods

2022-10-13 · Nils Feldhus, Leonhard Hennig, Maximilian Dustin Nasert, Christopher Ebert 외

Saliency maps can explain a neural model's predictions by identifying important input features. They are difficult to interpret for laypeople, especially for instances with many features. In order to make them more acces…

Abstractive Text SummarizationFeature ImportanceSentiment Analysistext-classification+2

A Psycholinguistic Evaluation of Language Models' Sensitivity to Argument Roles

2024-10-21 · Eun-Kyoung Rosa Lee, Sathvik Nair, Naomi Feldman

We present a systematic evaluation of large language models' sensitivity to argument roles, i.e., who did what to whom, by replicating psycholinguistic studies on human argument role processing. In three experiments, we …

SensitivitySentence

Comparing Apples to Oranges: LLM-powered Multimodal Intention Prediction in an Object Categorization Task

2024-04-12 · Hassan Ali, Philipp Allgeuer, Stefan Wermter

Human intention-based systems enable robots to perceive and interpret user actions to interact with humans and adapt to their behavior proactively. Therefore, intention prediction is pivotal in creating a natural interac…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Object Categorizationspeech-recognition+2

Towards Robust Evaluation of Visual Activity Recognition: Resolving Verb Ambiguity with Sense Clustering

2025-08-07 · Louie Hong Yao, Nicholas Jarvis, Tianyu Jiang arxiv

Evaluating visual activity recognition systems is challenging due to inherent ambiguities in verb semantics and image interpretation. When describing actions in images, synonymous verbs can refer to the same event (e.g.,…

Activity Recognition