paper-with-me

홈 › Papers

Question-Instructed Visual Descriptions for Zero-Shot Video Question Answering

2024-02-16 · David Romero, Thamar Solorio

We present Q-ViD, a simple approach for video question answering (video QA), that unlike prior methods, which are based on complex architectures, computationally expensive pipelines or use closed models like GPTs, Q-ViD relies on a single instruction-aware open vision-language model (InstructBLIP) to tackle videoQA using frame descriptions. Specifically, we create captioning instruction prompts that rely on the target questions about the videos and leverage InstructBLIP to obtain video frame captions that are useful to the task at hand. Subsequently, we form descriptions of the whole video using the question-dependent frame captions, and feed that information, along with a question-answering prompt, to a large language model (LLM). The LLM is our reasoning module, and performs the final step of multiple-choice QA. Our simple Q-ViD framework achieves competitive or even higher performances than current state of the art models on a diverse range of videoQA benchmarks, including NExT-QA, STAR, How2QA, TVQA and IntentQA.

📄 PDF Abstract BibTeX arXiv:2402.10698

Code (1)

daromog/q-vid 공식 구현 pytorch

Tasks

Language ModelingLanguage ModellingLarge Language ModelMultiple-choiceQuestion AnsweringVideo Question AnsweringZero-Shot Video Question Answer

Similar Papers 제목 키워드 기반

Robotic Programmer: Video Instructed Policy Code Generation for Robotic Manipulation

2025-01-08 · Senwei Xie, Hongyu Wang, Zhanqi Xiao, Ruiping Wang 외

Zero-shot generalization across various robots, tasks and environments remains a significant challenge in robotic manipulation. Policy code generation methods use executable code to connect high-level task descriptions a…

Code GenerationLanguage ModelingLanguage ModellingLarge Language Model+1

Rewriting Conversational Utterances with Instructed Large Language Models

2024-10-10 · Elnara Galimzhanova, Cristina Ioana Muntean, Franco Maria Nardini, Raffaele Perego 외

Many recent studies have shown the ability of large language models (LLMs) to achieve state-of-the-art performance on many NLP tasks, such as question answering, text summarization, coding, and translation. In some cases…

Conversational SearchQuestion AnsweringText Summarization

Zero-Shot Visual Reasoning by Vision-Language Models: Benchmarking and Analysis

2024-08-27 · Aishik Nagar, Shantanu Jaiswal, Cheston Tan

Vision-language models (VLMs) have shown impressive zero- and few-shot performance on real-world visual question answering (VQA) benchmarks, alluding to their capabilities as visual reasoning engines. However, the benchm…

BenchmarkingLarge Language ModelQuestion AnsweringVisual Question Answering+3

Evaluating Instruction-Tuned Large Language Models on Code Comprehension and Generation

2023-08-02 · Zhiqiang Yuan, Junwei Liu, Qiancheng Zi, Mingwei Liu 외

In this work, we evaluate 10 open-source instructed LLMs on four representative code comprehension and generation tasks. We have the following main findings. First, for the zero-shot setting, instructed LLMs are very com…

ZEST: Zero-shot Learning from Text Descriptions using Textual Similarity and Visual Summarization

2020-10-07 · Findings of the Association for Computational Linguistics 2020 · Tzuf Paz-Argaman, Yuval Atzmon, Gal Chechik, Reut Tsarfaty

We study the problem of recognizing visual entities from the textual descriptions of their classes. Specifically, given birds' images with free-text descriptions of their species, we learn to classify images of previousl…

Zero-Shot Learning