paper-with-me

홈 › Papers

Describing Common Human Visual Actions in Images

2015-06-07 · Matteo Ruggero Ronchi, Pietro Perona

Which common human actions and interactions are recognizable in monocular still images? Which involve objects and/or other people? How many is a person performing at a time? We address these questions by exploring the actions and interactions that are detectable in the images of the MS COCO dataset. We make two main contributions. First, a list of 140 common visual actions', obtained by analyzing the largest on-line verb lexicon currently available for English (VerbNet) and human sentences used to describe images in MS COCO. Second, a complete set of annotations for those visual actions', composed of subject-object and associated verb, which we call COCO-a (a for `actions'). COCO-a is larger than existing action datasets in terms of number of actions and instances of these actions, and is unique because it is data-driven, rather than experimenter-biased. Other unique features are that it is exhaustive, and that all subjects and objects are localized. A statistical analysis of the accuracy of our annotations and of each action, interaction and subject-object combination is provided.

📄 PDF Abstract BibTeX arXiv:1506.02203

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

HL Dataset: Visually-grounded Description of Scenes, Actions and Rationales

2023-02-23 · Michele Cafagna, Kees Van Deemter, Albert Gatt

Current captioning datasets focus on object-centric captions, describing the visible objects in the image, e.g. "people eating food in a park". Although these datasets are useful to evaluate the ability of Vision & Langu…

Common Sense ReasoningVocal Bursts Intensity Prediction

Human Action Adverb Recognition: ADHA Dataset and A Three-Stream Hybrid Model

2018-02-04 · Bo Pang, Kaiwen Zha, Cewu Lu

We introduce the first benchmark for a new problem --- recognizing human action adverbs (HAA): "Adverbs Describing Human Actions" (ADHA). This is the first step for computer vision to change over from pattern recognition…

Action RecognitionImage CaptioningTemporal Action Localization

Multi-modal Cooking Workflow Construction for Food Recipes

2020-08-20 · Liangming Pan, Jingjing Chen, Jianlong Wu, Shaoteng Liu 외

Understanding food recipe requires anticipating the implicit causal effects of cooking actions, such that the recipe can be converted into a graph describing the temporal workflow of the recipe. This is a non-trivial tas…

Common Sense ReasoningDecoder

Shaping Visual Representations with Language for Few-shot Classification

2019-11-06 · ACL 2020 6 · Jesse Mu, Percy Liang, Noah Goodman

By describing the features and abstractions of our world, language is a crucial tool for human learning and a promising source of supervision for machine learning models. We use language to improve few-shot visual classi…

ClassificationGeneral ClassificationMeta-LearningRepresentation Learning

Give Me Something to Eat: Referring Expression Comprehension with Commonsense Knowledge

2020-06-02 · Peng Wang, Dongyang Liu, Hui Li, Qi Wu

Conventional referring expression comprehension (REF) assumes people to query something from an image by describing its visual appearance and spatial location, but in practice, we often ask for an object by describing it…

16kReferring ExpressionReferring Expression Comprehension