Zero-Shot Video Question Answer
17개 벤치마크 · 논문 85편 · 이 태스크의 논문 보기 →
Benchmarks
MSRVTT-QA
EgoSchema (fullset)
ActivityNet-QA
MSVD-QA
NExT-QA
EgoSchema (subset)
TGIF-QA
IntentQA
Video-MME
NExT-GQA
TVQA
VNBench
Video-MME (w/o subs)
STAR Benchmark
MVBench
Most implemented
Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Mistral 7B
Flamingo: a Visual Language Model for Few-Shot Learning
Papers
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering
Video Question Answering (VQA) inherently relies on multimodal reasoning, integrating visual, temporal, and linguistic cues to achieve a deeper understanding of video content. However, many existing methods rely on feedi…
Caption GenerationEgoSchemaMultimodal ReasoningQuestion Answering+4Qwen2.5-Omni Technical Report
In this report, we present Qwen2.5-Omni, an end-to-end multimodal model designed to perceive diverse modalities, including text, images, audio, and video, while simultaneously generating text and natural speech responses…
Automatic Speech Recognition (ASR)GSM8KInstruction FollowingLarge Language Model+4Agentic Keyframe Search for Video Question Answering
Video question answering (VideoQA) enables machines to extract and comprehend key information from videos through natural language interaction, which is a critical step towards achieving intelligence. However, the demand…
EgoSchemaQuestion AnsweringVideo Question AnsweringVideo Understanding+2VideoMind: A Chain-of-LoRA Agent for Long Video Reasoning
Videos, with their unique temporal dimension, demand precise grounded understanding, where answers are directly linked to visual, interpretable evidence. Despite significant breakthroughs in reasoning capabilities within…
Grounded Video Question AnsweringQuestion AnsweringTemporal LocalizationVideo Question Answering+2BIMBA: Selective-Scan Compression for Long-Range Video Question Answering
Video Question Answering (VQA) in long videos poses the key challenge of extracting relevant information and modeling long-range dependencies from many redundant frames. The self-attention mechanism provides a general so…
Video Question AnsweringZero-Shot Video Question AnswerENTER: Event Based Interpretable Reasoning for VideoQA
In this paper, we present ENTER, an interpretable Video Question Answering (VideoQA) system based on event graphs. Event graphs convert videos into graphical representations, where video events form the nodes and event-e…
Code GenerationEgoSchemaQuestion AnsweringVideo Question Answering+1