paper-with-me

홈 › Papers

HL Dataset: Visually-grounded Description of Scenes, Actions and Rationales

2023-02-23 · Michele Cafagna, Kees Van Deemter, Albert Gatt

Current captioning datasets focus on object-centric captions, describing the visible objects in the image, e.g. "people eating food in a park". Although these datasets are useful to evaluate the ability of Vision & Language models to recognize and describe visual content, they do not support controlled experiments involving model testing or fine-tuning, with more high-level captions, which humans find easy and natural to produce. For example, people often describe images based on the type of scene they depict ('people at a holiday resort') and the actions they perform ('people having a picnic'). Such descriptions draw on personal experience and commonsense assumptions. We present the High-Level Dataset a dataset extending 14997 images from the COCO dataset, aligned with a new set of 134,973 human-annotated (high-level) captions collected along three axes: scenes, actions, and rationales. We further extend this dataset with confidence scores collected from an independent set of readers, as well as a set of narrative captions generated synthetically, by combining each of the three axes. We describe this dataset and analyse it extensively. We also present baseline results for the High-Level Captioning task.

📄 PDF Abstract BibTeX arXiv:2302.12189

Code (1)

michelecafagna26/hl-dataset 공식 구현

Tasks

Common Sense ReasoningVocal Bursts Intensity Prediction

Similar Papers 제목 키워드 기반

Enabling Robots to Draw and Tell: Towards Visually Grounded Multimodal Description Generation

2021-01-14 · Ting Han, Sina Zarrieß

Socially competent robots should be equipped with the ability to perceive the world that surrounds them and communicate about it in a human-like manner. Representative skills that exhibit such ability include generating …

TS-RGBD Dataset: a Novel Dataset for Theatre Scenes Description for People with Visual Impairments

2023-08-02 · Leyla Benhamida, Khadidja Delloul, Slimane Larabi

Computer vision was long a tool used for aiding visually impaired people to move around their environment and avoid obstacles and falls. Solutions are limited to either indoor or outdoor scenes, which limits the kind of …

Action RecognitionImage CaptioningTemporal Action Localization

Natural-language-driven Simulation Benchmark and Copilot for Efficient Production of Object Interactions in Virtual Road Scenes

2023-12-07 · Kairui Yang, Zihao Guo, Gengjie Lin, Haotian Dong 외

We advocate the idea of the natural-language-driven(NLD) simulation to efficiently produce the object interactions between multiple objects in the virtual road scenes, for teaching and testing the autonomous driving syst…

Autonomous DrivingObject

Learning language through pictures

2015-06-11 · IJCNLP 2015 7 · Grzegorz Chrupała, Ákos Kádár, Afra Alishahi

We propose Imaginet, a model of learning visually grounded representations of language from coupled textual and visual input. The model consists of two Gated Recurrent Unit networks with shared word embeddings, and uses …

SentenceWord Embeddings

Resolving References in Visually-Grounded Dialogue via Text Generation

2023-09-23 · SIGdial 2023 9 · Bram Willemsen, Livia Qian, Gabriel Skantze

Vision-language models (VLMs) have shown to be effective at image retrieval based on simple text queries, but text-image retrieval based on conversational input remains a challenge. Consequently, if we want to use VLMs f…

Image RetrievalLanguage ModelingLanguage ModellingLarge Language Model+2