Who’s Doing What: Joint Modeling of Names and Verbs for Simultaneous Face and Pose Annotation
Given a corpus of news items consisting of images accompanied by text captions, we want to find out `whos doing what, i.e. associate names and action verbs in the captions to the face and body pose of the persons in the images. We present a joint model for simultaneously solving the image-caption correspondences and learning visual appearance models for the face and pose classes occurring in the corpus. These models can then be used to recognize people and actions in novel images without captions. We demonstrate experimentally that our joint face and pose model solves the correspondence problem better than earlier models covering only the face, and that it can perform recognition of new uncaptioned images.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
What Are They Doing? Joint Audio-Speech Co-Reasoning
In audio and speech processing, tasks usually focus on either the audio or speech modality, even when both sounds and human speech are present in the same audio clip. Recent Auditory Large Language Models (ALLMs) have ma…
Economics in Nouns and Verbs
Standard economic theory uses mathematics as its main means of understanding, and this brings clarity of reasoning and logical power. But there is a drawback: algebraic mathematics restricts economic modeling to what can…
Grounded Video Situation Recognition
Dense video understanding requires answering several questions such as who is doing what to whom, with what, how, why, and where. Recently, Video Situation Recognition (VidSitu) is framed as a task for structured predict…
DescriptiveStructured PredictionVideo UnderstandingAdverbs in plWordNet: Theory and Implementation
Adverbs are seldom well represented in wordnets. Princeton WordNet, for example, derives from adjectives practically all its adverbs and whatever involvement they have. GermaNet stays away from this part of speech. Adver…
AllQuantifying Gender Bias Towards Politicians in Cross-Lingual Language Models
Recent research has demonstrated that large pre-trained language models reflect societal biases expressed in natural language. The present paper introduces a simple method for probing language models to conduct a multili…
Language ModelingLanguage ModellingProbing Language Models