paper-with-me

홈 › Papers

Towards Task Understanding in Visual Settings

2018-11-28 · Sebastin Santy, Wazeer Zulfikar, Rishabh Mehrotra, Emine Yilmaz

We consider the problem of understanding real world tasks depicted in visual images. While most existing image captioning methods excel in producing natural language descriptions of visual scenes involving human tasks, there is often the need for an understanding of the exact task being undertaken rather than a literal description of the scene. We leverage insights from real world task understanding systems, and propose a framework composed of convolutional neural networks, and an external hierarchical task ontology to produce task descriptions from input images. Detailed experiments highlight the efficacy of the extracted descriptions, which could potentially find their way in many applications, including image alt text generation.

📄 PDF Abstract BibTeX arXiv:1811.11833

Code (0)

등록된 구현이 없습니다.

Tasks

Image CaptioningText Generation

Similar Papers 제목 키워드 기반

TABLET: A Large-Scale Dataset for Robust Visual Table Understanding

2025-09-25 · Iñigo Alonso, Imanol Miranda, Eneko Agirre, Mirella Lapata arxiv

While table understanding increasingly relies on pixel-only settings, current benchmarks predominantly use synthetic renderings that lack the complexity and visual diversity of real-world tables. Additionally, existing v…

Grounded Agreement Games: Emphasizing Conversational Grounding in Visual Dialogue Settings

2019-08-29 · David Schlangen

Where early work on dialogue in Computational Linguistics put much emphasis on dialogue structure and its relation to the mental states of the dialogue participants (e.g., Allen 1979, Grosz & Sidner 1986), current work m…

ChatbotVisual Dialog

RoadscapesQA: A Multitask, Multimodal Dataset for Visual Question Answering on Indian Roads

2026-02-13 · Vijayasri Iyer, Maahin Rathinagiriswaran, Jyothikamalesh S arxiv

Understanding road scenes is essential for autonomous driving, as it enables systems to interpret visual surroundings to aid in effective decision-making. We present Roadscapes, a multitask multimodal dataset consisting …

Visual Question AnsweringScene UnderstandingAutonomous Driving

In-the-Wild Video Question Answering

2022-10-01 · COLING 2022 10 · Santiago Castro, Naihao Deng, Pingxuan Huang, Mihai Burzo 외

Existing video understanding datasets mostly focus on human interactions, with little attention being paid to the “in the wild” settings, where the videos are recorded outdoors. We propose WILDQA, a video understanding d…

Evidence SelectionQuestion AnsweringVideo Question AnsweringVideo Understanding

WildQA: In-the-Wild Video Question Answering

2022-09-14 · Santiago Castro, Naihao Deng, Pingxuan Huang, Mihai Burzo 외

Existing video understanding datasets mostly focus on human interactions, with little attention being paid to the "in the wild" settings, where the videos are recorded outdoors. We propose WILDQA, a video understanding d…

Evidence SelectionQuestion AnsweringVideo Question AnsweringVideo Understanding