paper-with-me

Zero-Shot Video Question Answer

17개 벤치마크 · 논문 85편 · 이 태스크의 논문 보기 →

Benchmarks

MSRVTT-QA

결과 60개

EgoSchema (fullset)

결과 58개

ActivityNet-QA

결과 56개

MSVD-QA

결과 56개

NExT-QA

결과 54개

EgoSchema (subset)

결과 28개

TGIF-QA

결과 28개

IntentQA

결과 26개

Video-MME

결과 22개

NExT-GQA

결과 18개

TVQA

결과 18개

VNBench

결과 18개

Video-MME (w/o subs)

결과 18개

STAR Benchmark

결과 8개

MVBench

결과 4개

Most implemented

Mistral 7B

2023-10-10 · 구현 6개

Papers

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering

2025-04-25 · Noriyuki Kugo, Xiang Li, Zixin Li, Ashish Gupta 외

Video Question Answering (VQA) inherently relies on multimodal reasoning, integrating visual, temporal, and linguistic cues to achieve a deeper understanding of video content. However, many existing methods rely on feedi…

Caption GenerationEgoSchemaMultimodal ReasoningQuestion Answering+4

Qwen2.5-Omni Technical Report

2025-03-26 · Jin Xu, Zhifang Guo, Jinzheng He, Hangrui Hu 외

In this report, we present Qwen2.5-Omni, an end-to-end multimodal model designed to perceive diverse modalities, including text, images, audio, and video, while simultaneously generating text and natural speech responses…

Automatic Speech Recognition (ASR)GSM8KInstruction FollowingLarge Language Model+4

Agentic Keyframe Search for Video Question Answering

2025-03-20 · Sunqi Fan, Meng-Hao Guo, Shuojin Yang

Video question answering (VideoQA) enables machines to extract and comprehend key information from videos through natural language interaction, which is a critical step towards achieving intelligence. However, the demand…

EgoSchemaQuestion AnsweringVideo Question AnsweringVideo Understanding+2

VideoMind: A Chain-of-LoRA Agent for Long Video Reasoning

2025-03-17 · Ye Liu, Kevin Qinghong Lin, Chang Wen Chen, Mike Zheng Shou

Videos, with their unique temporal dimension, demand precise grounded understanding, where answers are directly linked to visual, interpretable evidence. Despite significant breakthroughs in reasoning capabilities within…

Grounded Video Question AnsweringQuestion AnsweringTemporal LocalizationVideo Question Answering+2

BIMBA: Selective-Scan Compression for Long-Range Video Question Answering

2025-03-12 · CVPR 2025 1 · Md Mohaiminul Islam, Tushar Nagarajan, Huiyu Wang, Gedas Bertasius 외

Video Question Answering (VQA) in long videos poses the key challenge of extracting relevant information and modeling long-range dependencies from many redundant frames. The self-attention mechanism provides a general so…

Video Question AnsweringZero-Shot Video Question Answer

ENTER: Event Based Interpretable Reasoning for VideoQA

2025-01-24 · Hammad Ayyubi, Junzhang Liu, Ali Asgarov, Zaber Ibn Abdul Hakim 외

In this paper, we present ENTER, an interpretable Video Question Answering (VideoQA) system based on event graphs. Event graphs convert videos into graphical representations, where video events form the nodes and event-e…

Code GenerationEgoSchemaQuestion AnsweringVideo Question Answering+1

전체 85편 보기 →