Fill-in-the-Blank: A Challenging Video Understanding Evaluation Framework
We propose fill-in-the-blanks as a video understanding evaluation framework. The task tests a model's understanding of a video by requiring the model to predict a masked noun phrase in the caption of the video, given the video and the surrounding text. To this end, we introduce a novel dataset consisting of 28,000 videos and fill-in-the-blank tests with multiple correct answers. The task and the dataset are challenging for the current state-of-the-art systems to solve. This task also does not share the weaknesses of the current state of the art language-informed video understanding tasks, namely: (1) video question answering using multiple-choice questions, where models perform relatively well because they exploit linguistic biases in the task formulation; and (2) video captioning, which relies on an open-ended evaluation framework that is often inaccurate because system answers may be perceived as incorrect if they differ in form from the ground truth.
Code (0)
등록된 구현이 없습니다.
Tasks
Multiple-choiceQuestion AnsweringVideo CaptioningVideo Question AnsweringVideo UnderstandingSimilar Papers 제목 키워드 기반
FIBER: Fill-in-the-Blanks as a Challenging Video Understanding Evaluation Framework
We propose fill-in-the-blanks as a video understanding evaluation framework and introduce FIBER -- a novel dataset consisting of 28,000 videos and descriptions in support of this evaluation framework. The fill-in-the-bla…
Language ModellingMultiple-choiceQuestion AnsweringVideo Captioning+2Video Fill in the Blank with Merging LSTMs
Given a video and its incomplete textural description with missing words, the Video-Fill-in-the-Blank (ViFitB) task is to automatically find the missing word. The contextual information of the sentences are important to …
A dataset and exploration of models for understanding video data through fill-in-the-blank question-answering
While deep convolutional neural networks frequently approach or exceed human-level performance at benchmark tasks involving static images, extending this success to moving images is not straightforward. Having models whi…
DescriptiveLanguage ModelingLanguage Modellingobject-detection+2Video Fill In the Blank using LR/RL LSTMs with Spatial-Temporal Attentions
Given a video and a description sentence with one missing word (we call it the "source sentence"), Video-Fill-In-the-Blank (VFIB) problem is to find the missing word automatically. The contextual information of the sente…
SentenceQUIET: A Multi-Blank Cascaded Story Cloze Benchmark for LLM Creative Generation Capability
Large language models (LLMs) face a dual challenge in creative capability evaluation: existing benchmarks (e.g., Story Cloze Test, HellaSwag) measure models' discriminative ability over narrative continuation using multi…
Logical ReasoningCloze Test