paper-with-me

홈 › Papers

Seeing What You're Told: Sentence-Guided Activity Recognition In Video

2013-08-19 · CVPR 2014 6 · N. Siddharth, Andrei Barbu, Jeffrey Mark Siskind

We present a system that demonstrates how the compositional structure of events, in concert with the compositional structure of language, can interplay with the underlying focusing mechanisms in video action recognition, thereby providing a medium, not only for top-down and bottom-up integration, but also for multi-modal integration between vision and language. We show how the roles played by participants (nouns), their characteristics (adjectives), the actions performed (verbs), the manner of such actions (adverbs), and changing spatial relations between participants (prepositions) in the form of whole sentential descriptions mediated by a grammar, guides the activity-recognition process. Further, the utility and expressiveness of our framework is demonstrated by performing three separate tasks in the domain of multi-activity videos: sentence-guided focus of attention, generation of sentential descriptions of video, and query-based video search, simply by leveraging the framework in different manners.

📄 PDF Abstract BibTeX arXiv:1308.4189

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionActivity RecognitionSentenceTemporal Action Localization

Similar Papers 제목 키워드 기반

Tell Me More: A Dataset of Visual Scene Description Sequences

2019-10-01 · WS 2019 10 · Nikolai Ilinykh, Sina Zarrie{\ss}, David Schlangen

We present a dataset consisting of what we call image description sequences, which are multi-sentence descriptions of the contents of an image. These descriptions were collected in a pseudo-interactive setting, where the…

Image DescriptionSentence

Narrative Interpolation for Generating and Understanding Stories

2020-08-17 · Su Wang, Greg Durrett, Katrin Erk

We propose a method for controlled narrative/story generation where we are able to guide the model to produce coherent narratives with user-specified target endings by interpolation: for example, we are told that Jim wen…

SentenceStory Generation

Learning to Perform Role-Filler Binding with Schematic Knowledge

2019-02-24 · Catherine Chen, Qihong Lu, Andre Beukers, Christopher Baldassano 외

Through specific experiences, humans learn relationships underlying the structure of events in the world. Schema theory suggests that we organize this information in mental frameworks called "schemata," which represent o…

Question AnsweringSentence

Effective Explanations Support Planning Under Uncertainty

2026-05-08 · Hanqi Zhou, Britt Besch, Charley M. Wu, Tobias Gerstenberg arxiv

Explaining how to get from A to B can be challenging. It requires mentally simulating what the listener will do based on what they are told. To capture this process, we propose a computational model that converts utteran…

Prompt-Guided Generation of Structured Chest X-Ray Report Using a Pre-trained LLM

2024-04-17 · Hongzhao Li, Hongyu Wang, Xia Sun, Hua He 외

Medical report generation automates radiology descriptions from images, easing the burden on physicians and minimizing errors. However, current methods lack structured outputs and physician interactivity for clear, clini…

AnatomyLanguage ModelingLanguage ModellingLarge Language Model+2