paper-with-me

Papers

Grounding Semantic Roles in Images

2018-10-01 · EMNLP 2018 10 · Carina Silberer, Manfred Pinkal

We address the task of visual semantic role labeling (vSRL), the identification of the participants of a situation or event in a visual scene, and their labeling with their semantic relations to the event or situation. We render candidate participants as image regions of objects, and train a model which learns to ground roles in the regions which depict the corresponding participant. Experimental results demonstrate that we can train a vSRL model without reliance on prohibitive image-based role annotations, by utilizing noisy data which we extract automatically from image captions using a linguistic SRL system. Furthermore, our model induces frame{---}semantic visual representations, and their comparison to previous work on supervised visual verb sense disambiguation yields overall better results.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Image CaptioningQuestion AnsweringSemantic Role Labeling

Similar Papers 제목 키워드 기반

Grounded Situation Recognition

2020-03-26 · ECCV 2020 8 · Sarah Pratt, Mark Yatskar, Luca Weihs, Ali Farhadi 외

We introduce Grounded Situation Recognition (GSR), a task that requires producing structured semantic summaries of images describing: the primary activity, entities engaged in the activity with their roles (e.g. agent, t…

Grounded Situation RecognitionImage RetrievalRetrievalSituation Recognition

Video Object Grounding using Semantic Roles in Language Description

2020-03-24 · CVPR 2020 6 · Arka Sadhu, Kan Chen, Ram Nevatia

We explore the task of Video Object Grounding (VOG), which grounds objects in videos referred to in natural language descriptions. Previous methods apply image grounding based algorithms to address VOG, fail to explore t…

ObjectPosition

GSRFormer: Grounded Situation Recognition Transformer with Alternate Semantic Attention Refinement

2022-08-18 · Zhi-Qi Cheng, Qi Dai, SiYao Li, Teruko Mitamura 외

Grounded Situation Recognition (GSR) aims to generate structured semantic summaries of images for "human-like" event understanding. Specifically, GSR task not only detects the salient activity verb (e.g. buying), but als…

Grounded Situation RecognitionImage Captioningobject-detectionObject Detection+1

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition

2026-04-25 · Balaji Darur, Amanmeet Garg, Makarand Tapaswi arxiv

Video Situation Recognition (VidSitu) addresses the challenging problem of "who did what to whom, with what, how, and where" in a video. It tests thorough video understanding by requiring identification of salient action…

Situation RecognitionVisual Grounding

Roles of MLLMs in Visually Rich Document Retrieval for RAG: A Survey

2025-12-16 · Xiantao Zhang arxiv

Visually rich documents (VRDs) challenge retrieval-augmented generation (RAG) with layout-dependent semantics, brittle OCR, and evidence spread across complex figures and structured tables. This survey examines how Multi…