paper-with-me

Papers MVBench

“MVBench” 태그가 달린 논문 19편 · 필터 해제

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

2025-06-30 · Haoji Zhang, Yiqin Wang, Yansong Tang, Yong liu 외

Benefiting from the advances in large language models and cross-modal alignment, existing multimodal large language models have achieved prominent performance in image and short video understanding. However, the understa…

cross-modal alignmentEgoSchemaMMEMVBench+2

GAM-Agent: Game-Theoretic and Uncertainty-Aware Collaboration for Complex Visual Reasoning

2025-05-29 · Jusheng Zhang, Yijia Fan, Wenjun Lin, Ruiqi Chen 외

We propose GAM-Agent, a game-theoretic multi-agent framework for enhancing vision-language reasoning. Unlike prior single-agent or monolithic models, GAM-Agent formulates the reasoning process as a non-zero-sum game betw…

Multimodal ReasoningMVBenchVisual Reasoning

VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding

2025-05-02 · Zongxia Li, Xiyang Wu, Guangyao Shi, Yubin Qin 외

Synthetic video generation has gained significant attention for its realism and broad applications, but remains prone to violations of common sense and physical laws. This highlights the need for reliable abnormality det…

Anomaly DetectionCommon Sense ReasoningHallucinationMVBench+2

VideoPASTA: 7K Preference Pairs That Matter for Video-LLM Alignment

2025-04-18 · Yogesh Kulkarni, Pooyan Fazli

Video-language models (Video-LLMs) excel at understanding video content but struggle with spatial relationships, temporal ordering, and cross-frame continuity. To address these limitations, we introduce VideoPASTA (Prefe…

MVBench

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

2025-04-09 · Xinhao Li, Ziang Yan, Desen Meng, Lu Dong 외

Recent advancements in reinforcement learning have significantly advanced the reasoning capabilities of multimodal large language models (MLLMs). While approaches such as Group Relative Policy Optimization (GRPO) and rul…

MVBenchObject TrackingVideo Understanding

Mobile-VideoGPT: Fast and Accurate Video Understanding Language Model

2025-03-27 · Abdelrahman Shaker, Muhammad Maaz, Chenhui Gou, Hamid Rezatofighi 외

Video understanding models often struggle with high computational requirements, extensive parameter counts, and slow inference speed, making them inefficient for practical use. To tackle these challenges, we propose Mobi…

EgoSchemaLanguage ModelingLanguage ModellingMVBench+2

Video-R1: Reinforcing Video Reasoning in MLLMs

2025-03-27 · Kaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo 외

Inspired by DeepSeek-R1's success in eliciting reasoning abilities through rule-based reinforcement learning (RL), we introduce Video-R1 as the first attempt to systematically explore the R1 paradigm for incentivizing vi…

MVBenchReinforcement Learning (RL)Spatial Reasoning

LLaVAction: evaluating and training multi-modal large language models for action recognition

2025-03-24 · Shaokai Ye, Haozhe Qi, Alexander Mathis, Mackenzie W. Mathis

Understanding human behavior requires measuring behavioral actions. Due to its complexity, behavior is best mapped onto a rich, semantic structure such as language. The recent development of multi-modal large language mo…

Action RecognitionAction UnderstandingEgoSchemaMVBench+1

Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition

2024-12-12 · Zhisheng Zhong, Chengyao Wang, Yuqi Liu, Senqiao Yang 외

As Multi-modal Large Language Models (MLLMs) evolve, expanding beyond single-domain capabilities is essential to meet the demands for more versatile and efficient AI. However, previous omni-models have insufficiently exp…

EgoSchemaMMEMM-Vet+5

VideoSAVi: Self-Aligned Video Language Models without Human Supervision

2024-12-01 · Yogesh Kulkarni, Pooyan Fazli

Recent advances in video-large language models (Video-LLMs) have led to significant progress in video understanding. Current preference optimization methods often rely on proprietary APIs or ground-truth captions to gene…

EgoSchemaMVBenchSpatial ReasoningVideo Understanding

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models

2024-11-17 · Tingyu Qu, Mingxiao Li, Tinne Tuytelaars, Marie-Francine Moens

Recent advances in multimodal Large Language Models (LLMs) have shown great success in understanding multi-modal contents. For video understanding tasks, training-based video LLMs are difficult to build due to the scarci…

MVBenchVideo-based Generative Performance BenchmarkingVideo-based Generative Performance Benchmarking (Consistency)Video-based Generative Performance Benchmarking (Contextual Understanding)+5

Enhancing Temporal Modeling of Video LLMs via Time Gating

2024-10-08 · Zi-Yuan Hu, Yiwu Zhong, Shijia Huang, Michael R. Lyu 외

Video Large Language Models (Video LLMs) have achieved impressive performance on video-and-language tasks, such as video question answering. However, most existing Video LLMs neglect temporal information in video data, l…

MVBenchQuestion AnsweringVideo Question AnsweringVideo Understanding

VideoLLaMB: Long-context Video Understanding with Recurrent Memory Bridges

2024-09-02 · Yuxuan Wang, Cihang Xie, Yang Liu, Zilong Zheng

Recent advancements in large-scale video-language models have shown significant potential for real-time planning and detailed interactions. However, their high computational demands and the scarcity of annotated datasets…

GPUMVBenchVideo Understanding

CogVLM2: Visual Language Models for Image and Video Understanding

2024-08-29 · Wenyi Hong, Weihan Wang, Ming Ding, Wenmeng Yu 외

Beginning with VisualGLM and CogVLM, we are continuously exploring VLMs in pursuit of enhanced vision-language fusion, efficient higher-resolution architecture, and broader modalities and applications. Here we propose th…

MM-VetMVBenchTextVQAVideo Understanding+1

Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

2024-08-26 · Jiajun Fei, Dian Li, Zhidong Deng, Zekun Wang 외

Multi-modal large language models (MLLMs) have demonstrated considerable potential across various downstream tasks that require cross-domain knowledge. MLLMs capable of processing videos, known as Video-MLLMs, have attra…

Large Language ModelMVBenchVideo Understanding

VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding

2024-06-13 · Muhammad Maaz, Hanoona Rasheed, Salman Khan, Fahad Khan

Building on the advances of language models, Large Multimodal Models (LMMs) have contributed significant improvements in video understanding. While the current video LMMs utilize advanced Large Language Models (LLMs), th…

Dense Video CaptioningMVBenchQuestion AnsweringVCGBench-Diverse+10

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

2024-04-25 · arXiv 2024 4 · Lin Xu, Yilin Zhao, Daquan Zhou, Zhijie Lin 외

Vision-language pre-training has significantly elevated performance across a wide range of image-language applications. Yet, the pre-training process for video-related tasks demands exceptionally large computational and …

Dense CaptioningMVBenchVideo-based Generative Performance BenchmarkingVideo-based Generative Performance Benchmarking (Consistency)+7

ST-LLM: Large Language Models Are Effective Temporal Learners

2024-03-30 · Ruyang Liu, Chen Li, Haoran Tang, Yixiao Ge 외

Large Language Models (LLMs) have showcased impressive capabilities in text comprehension and generation, prompting research efforts towards video LLMs to facilitate human-AI interaction at the video level. However, how …

MVBenchReading ComprehensionVideo-based Generative Performance BenchmarkingVideo Question Answering+1

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

2023-11-28 · CVPR 2024 1 · Kunchang Li, Yali Wang, Yinan He, Yizhuo Li 외

With the rapid development of Multi-modal Large Language Models (MLLMs), a number of diagnostic benchmarks have recently emerged to evaluate the comprehension capabilities of these models. However, most benchmarks predom…

3D Question Answering (3D-QA)DiagnosticFairnessMultiple-choice+12
1–19 / 19