paper-with-me

MSRVTT-QA

홈페이지 · 논문 66편

The MSR-VTT-QA dataset is a benchmark for the task of Visual Question Answering (VQA) on the MSR-VTT (Microsoft Research Video to Text) dataset. The MSR-VTT-QA benchmark is used to evaluate models on their ability to answer questions based on these videos. It's part of the tasks that this dataset is used for, along with Video Retrieval, Video Captioning, Zero-Shot Video Question Answering, Zero-Shot Video Retrieval, and Text-to-Video Generation.

벤치마크

Zero-Shot Video Question Answer on MSRVTT-QA 결과 60개
Visual Question Answering (VQA) on MSRVTT-QA 결과 34개
Video Question Answering on MSRVTT-QA 결과 14개
Visual Question Answering on MSRVTT-QA 결과 8개
Zero-Shot Learning on MSRVTT-QA 결과 1개