Video Question Answering
28개 벤치마크 · 논문 620편 · 이 태스크의 논문 보기 →
Benchmarks
NExT-QA
ActivityNet-QA
TVBench
MVBench
STAR Benchmark
OVBench
MSRVTT-QA
AGQA 2.0 balanced
How2QA
MSRVTT-MC
iVQA
IntentQA
Perception Test
SUTD-TrafficQA
TVQA
WildQA
LSMDC-MC
NExT-QA (Efficient)
RoadTextVQA
DramaQA
Howto100M-QA
LSMDC-FiB
MSR-VTT
MSR-VTT-MC
MSVD-QA
TGIF-QA
VLEP
VideoQA
Most implemented
Is Space-Time Attention All You Need for Video Understanding?
Visual Instruction Tuning
Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Flamingo: a Visual Language Model for Few-Shot Learning
Papers
Can We Trust Video Hallucination Detectors? VidHalLoc for Evaluating the Evaluators
Video-language models and video agents can produce hallucinations that conflict with spatiotemporal evidence. Existing benchmarks mainly evaluate model hallucinations, and heterogeneous mechanisms make detector reliabili…
Video Question AnsweringVideo CaptioningEgoSIS: From Factorized Visual Ego-Transitions to Motion-Canonical Spatial Evidence for UAV Reasoning
UAV video question answering requires separating camera motion from changes in the scene, but RGB-only multimodal models receive no explicit, stable reference for that separation. We present EgoSIS, a pose-free adapter t…
Video Question AnsweringSpatial ReasoningEM^2Mem: Event-Centric Multimodal Memory for Large Language Models
Multimodal memory offers a scalable interface for long-video question answering, but existing methods often retrieve captions, frames, transcripts, summaries, or graph facts as isolated fragments. Although searchable, su…
Video Question AnsweringPost-Training VLMs for Video Mistake Detection
Human mistakes are inevitable when following instructions, yet they can lead to severe consequences. As such, there has been an increased interest in developing methods for detecting mistakes in videos, with current meth…
Video Question AnsweringVideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding
Video understanding is a fundamental task for evaluating the capabilities of multimodal large language models (MLLMs). However, existing leading models have already achieved approximately 90% accuracy on the Video-MME le…
Video Question AnsweringPhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs
Embodied intelligence and world models require video understanding systems to go beyond recognizing objects and actions and develop an understanding of physical regularities. However, despite their strong performance on …
Video Question Answering