paper-with-me

Video Question Answering

28개 벤치마크 · 논문 620편 · 이 태스크의 논문 보기 →

Benchmarks

NExT-QA

결과 48개

ActivityNet-QA

결과 36개

TVBench

결과 28개

MVBench

결과 22개

STAR Benchmark

결과 17개

OVBench

결과 16개

MSRVTT-QA

결과 14개

AGQA 2.0 balanced

결과 8개

How2QA

결과 8개

MSRVTT-MC

결과 7개

iVQA

결과 7개

IntentQA

결과 6개

Perception Test

결과 6개

SUTD-TrafficQA

결과 6개

TVQA

결과 6개

WildQA

결과 5개

LSMDC-MC

결과 2개

NExT-QA (Efficient)

결과 2개

RoadTextVQA

결과 2개

DramaQA

결과 1개

Howto100M-QA

결과 1개

LSMDC-FiB

결과 1개

MSR-VTT

결과 1개

MSR-VTT-MC

결과 1개

MSVD-QA

결과 1개

TGIF-QA

결과 1개

VLEP

결과 1개

VideoQA

결과 1개

Most implemented

Visual Instruction Tuning

2023-04-17 · 구현 13개

Papers

Can We Trust Video Hallucination Detectors? VidHalLoc for Evaluating the Evaluators

2026-09-09 · Xinyu Chen, Adnan Mahmood, Mark Dras arxiv

Video-language models and video agents can produce hallucinations that conflict with spatiotemporal evidence. Existing benchmarks mainly evaluate model hallucinations, and heterogeneous mechanisms make detector reliabili…

Video Question AnsweringVideo Captioning

EgoSIS: From Factorized Visual Ego-Transitions to Motion-Canonical Spatial Evidence for UAV Reasoning

2026-09-08 · Jingpu Yang, Fengxian Ji, Mingxuan Cui, Yilin Sun 외 arxiv

UAV video question answering requires separating camera motion from changes in the scene, but RGB-only multimodal models receive no explicit, stable reference for that separation. We present EgoSIS, a pose-free adapter t…

Video Question AnsweringSpatial Reasoning

EM^2Mem: Event-Centric Multimodal Memory for Large Language Models

2026-09-01 · Yijun Chen, Yaqi Zheng, Yanya Li, Boyi Xiao 외 hf

Multimodal memory offers a scalable interface for long-video question answering, but existing methods often retrieve captions, frames, transcripts, summaries, or graph facts as isolated fragments. Although searchable, su…

Video Question Answering

Post-Training VLMs for Video Mistake Detection

2026-08-28 · Federico Spurio, Olga Zatsarynna, Lars Doorenbos, Emad Bahrami 외 arxiv

Human mistakes are inevitable when following instructions, yet they can lead to severe consequences. As such, there has been an increased interest in developing methods for detecting mistakes in videos, with current meth…

Video Question Answering

VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding

2026-08-12 · Fan Zhang, Guangming Yao, Jinyang Wu, Hao Wu 외 hf

Video understanding is a fundamental task for evaluating the capabilities of multimodal large language models (MLLMs). However, existing leading models have already achieved approximately 90% accuracy on the Video-MME le…

Video Question Answering

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs

2026-08-03 · Zhongjie Ba, Shengwang Xu, Peng Cheng, Jinyang Zou 외 arxiv

Embodied intelligence and world models require video understanding systems to go beyond recognizing objects and actions and develop an understanding of physical regularities. However, despite their strong performance on …

Video Question Answering

전체 620편 보기 →