paper-with-me

Papers Zero-Shot Video Question Answer

“Zero-Shot Video Question Answer” 태그가 달린 논문 85편 · 필터 해제

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering

2025-04-25 · Noriyuki Kugo, Xiang Li, Zixin Li, Ashish Gupta 외

Video Question Answering (VQA) inherently relies on multimodal reasoning, integrating visual, temporal, and linguistic cues to achieve a deeper understanding of video content. However, many existing methods rely on feedi…

Caption GenerationEgoSchemaMultimodal ReasoningQuestion Answering+4

Qwen2.5-Omni Technical Report

2025-03-26 · Jin Xu, Zhifang Guo, Jinzheng He, Hangrui Hu 외

In this report, we present Qwen2.5-Omni, an end-to-end multimodal model designed to perceive diverse modalities, including text, images, audio, and video, while simultaneously generating text and natural speech responses…

Automatic Speech Recognition (ASR)GSM8KInstruction FollowingLarge Language Model+4

Agentic Keyframe Search for Video Question Answering

2025-03-20 · Sunqi Fan, Meng-Hao Guo, Shuojin Yang

Video question answering (VideoQA) enables machines to extract and comprehend key information from videos through natural language interaction, which is a critical step towards achieving intelligence. However, the demand…

EgoSchemaQuestion AnsweringVideo Question AnsweringVideo Understanding+2

VideoMind: A Chain-of-LoRA Agent for Long Video Reasoning

2025-03-17 · Ye Liu, Kevin Qinghong Lin, Chang Wen Chen, Mike Zheng Shou

Videos, with their unique temporal dimension, demand precise grounded understanding, where answers are directly linked to visual, interpretable evidence. Despite significant breakthroughs in reasoning capabilities within…

Grounded Video Question AnsweringQuestion AnsweringTemporal LocalizationVideo Question Answering+2

BIMBA: Selective-Scan Compression for Long-Range Video Question Answering

2025-03-12 · CVPR 2025 1 · Md Mohaiminul Islam, Tushar Nagarajan, Huiyu Wang, Gedas Bertasius 외

Video Question Answering (VQA) in long videos poses the key challenge of extracting relevant information and modeling long-range dependencies from many redundant frames. The self-attention mechanism provides a general so…

Video Question AnsweringZero-Shot Video Question Answer

ENTER: Event Based Interpretable Reasoning for VideoQA

2025-01-24 · Hammad Ayyubi, Junzhang Liu, Ali Asgarov, Zaber Ibn Abdul Hakim 외

In this paper, we present ENTER, an interpretable Video Question Answering (VideoQA) system based on event graphs. Event graphs convert videos into graphical representations, where video events form the nodes and event-e…

Code GenerationEgoSchemaQuestion AnsweringVideo Question Answering+1

LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

2025-01-07 · Shaolei Zhang, Qingkai Fang, Zhe Yang, Yang Feng

The advent of real-time large multimodal models (LMMs) like GPT-4o has sparked considerable interest in efficient LMMs. LMM frameworks typically encode visual inputs into vision tokens (continuous representations) and in…

GPUVisual Question Answering (VQA)Zero-Shot Video Question Answer

VidCtx: Context-aware Video Question Answering with Image Models

2024-12-23 · Andreas Goulas, Vasileios Mezaris, Ioannis Patras

To address computational and memory limitations of Large Multimodal Models in the Video Question-Answering task, several recent methods extract textual representations per frame (e.g., by captioning) and feed them to a L…

Large Language ModelQuestion AnsweringVideo Question AnsweringZero-Shot Video Question Answer

LinVT: Empower Your Image-level Large Language Model to Understand Videos

2024-12-06 · Lishuai Gao, Yujie Zhong, Yingsen Zeng, Haoxian Tan 외

Large Language Models (LLMs) have been widely used in various tasks, motivating us to develop an LLM-based assistant for videos. Instead of training from scratch, we propose a module to transform arbitrary well-trained i…

Language ModelingLanguage ModellingLarge Language ModelVideo Question Answering+3

Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension

2024-11-20 · Yongdong Luo, Xiawu Zheng, Xiao Yang, Guilin Li 외

Existing large video-language models (LVLMs) struggle to comprehend long videos correctly due to limited context. To address this problem, fine-tuning long-context LVLMs and employing GPT-based agents have emerged as pro…

GPUMMEobject-detectionObject Detection+9

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models

2024-11-17 · Tingyu Qu, Mingxiao Li, Tinne Tuytelaars, Marie-Francine Moens

Recent advances in multimodal Large Language Models (LLMs) have shown great success in understanding multi-modal contents. For video understanding tasks, training-based video LLMs are difficult to build due to the scarci…

MVBenchVideo-based Generative Performance BenchmarkingVideo-based Generative Performance Benchmarking (Consistency)Video-based Generative Performance Benchmarking (Contextual Understanding)+5

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance

2024-11-04 · Ruyang Liu, Haoran Tang, Haibo Liu, Yixiao Ge 외

The past year has witnessed the significant advancement of video-based large language models. However, the challenge of developing a unified model for both short and long video understanding remains unresolved. Most exis…

Caption GenerationMultiple-choiceVideo-based Generative Performance BenchmarkingVideo-based Generative Performance Benchmarking (Consistency)+7

TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning

2024-10-25 · Xiangyu Zeng, Kunchang Li, Chenting Wang, Xinhao Li 외

Multimodal Large Language Models (MLLMs) have demonstrated impressive performance in short video understanding. However, understanding long-form videos still remains challenging for MLLMs. This paper proposes TimeSuite, …

EgoSchemaHallucinationHighlight DetectionMoment Retrieval+3

GPT-4o System Card

2024-10-25 · OpenAI, :, Aaron Hurst, Adam Lerer 외

GPT-4o is an autoregressive omni model that accepts as input any combination of text, audio, image, and video, and generates any combination of text, audio, and image outputs. It's trained end-to-end across text, vision,…

Multiple-choiceSpatial ReasoningVideo Question AnsweringVisual Question Answering (VQA)+1

LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

2024-10-22 · Xiaoqian Shen, Yunyang Xiong, Changsheng Zhao, Lemeng Wu 외

Multimodal Large Language Models (MLLMs) have shown promising progress in understanding and analyzing video content. However, processing long videos remains a significant challenge constrained by LLM's context size. To a…

Token ReductionVideo Question AnsweringVideo UnderstandingZero-Shot Video Question Answer

Video Instruction Tuning With Synthetic Data

2024-10-03 · Yuanhan Zhang, Jinming Wu, Wei Li, Bo Li 외

The development of video large multimodal models (LMMs) has been hindered by the difficulty of curating large amounts of high-quality raw data from the web. To address this, we propose an alternative approach by creating…

3D Question Answering (3D-QA)Instruction FollowingMultiple-choice+5

Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

2024-09-18 · Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang 외

We present the Qwen2-VL Series, an advanced upgrade of the previous Qwen-VL models that redefines the conventional predetermined-resolution approach in visual processing. Qwen2-VL introduces the Naive Dynamic Resolution …

Natural Language Visual GroundingTemporal Relation ExtractionVideo Question Answering+3

Question-Answering Dense Video Events

2024-09-06 · Hangyu Qin, Junbin Xiao, Angela Yao

This paper presents question-answering on dense video events, a novel task that answers and grounds dense-event questions in long videos, thus challenging MLLMs to faithfully comprehend and reason about multiple events o…

BenchmarkingQuestion AnsweringZero-Shot Video Question Answer

LLaVA-OneVision: Easy Visual Task Transfer

2024-08-06 · Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang 외

We present LLaVA-OneVision, a family of open large multimodal models (LMMs) developed by consolidating our insights into data, models, and visual representations in the LLaVA-NeXT blog series. Our experimental results de…

3D Question Answering (3D-QA)Multiple-choiceTemporal Relation Extraction+6

MiniCPM-V: A GPT-4V Level MLLM on Your Phone

2024-08-03 · Yuan YAO, Tianyu Yu, Ao Zhang, Chongyi Wang 외

The recent surge of Multimodal Large Language Models (MLLMs) has fundamentally reshaped the landscape of AI research and industry, shedding light on a promising path toward the next AI milestone. However, significant cha…

HallucinationMultiple-choiceOptical Character Recognition (OCR)Temporal Relation Extraction+1
1–20 / 85 다음 →