paper-with-me

Papers

VideoChat: Chat-Centric Video Understanding

2023-05-10 · Kunchang Li, Yinan He, Yi Wang, Yizhuo Li, Wenhai Wang, Ping Luo, Yali Wang, LiMin Wang, Yu Qiao

In this paper, we initiate an attempt of developing an end-to-end chat-centric video understanding system, coined as VideoChat. It integrates video foundation models and large language models via a learnable neural interface, excelling in spatiotemporal reasoning, event localization, and causal relationship inference. To instructively tune this system, we build a video-centric instruction dataset, composed of thousands of videos associated with detailed descriptions and conversations. This dataset emphasizes spatiotemporal reasoning and captures causal relationships, providing a valuable asset for training our chat-centric video understanding system. Preliminary qualitative experiments demonstrate the potential of our system across a broad spectrum of video applications, which could serve as a simple prototype system for future research on chat-centric video understanding. Access our code and data at https://github.com/OpenGVLab/Ask-Anything

📄 PDF Abstract BibTeX arXiv:2305.06355

Code (1)

opengvlab/ask-anything 공식 구현 pytorch

Tasks

Question AnsweringVideo-based Generative Performance BenchmarkingVideo-based Generative Performance Benchmarking (Consistency)Video-based Generative Performance Benchmarking (Contextual Understanding)Video-based Generative Performance Benchmarking (Correctness of Information)Video-based Generative Performance Benchmarking (Detail Orientation))Video-based Generative Performance Benchmarking (Temporal Understanding)Video Question AnsweringVideo UnderstandingZero-Shot Video Question Answer

Similar Papers 제목 키워드 기반

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

2026-07-16 · Xinhao Li, Yuhan Zhu, Xiangyu Zeng, Yuhao Dong 외 arxiv

Recent advances in video understanding have spanned motion, long video, and streaming interaction, driving this field toward real-world applications. Despite this progress, current open-source models remain limited in se…

Computational Efficiency

TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning

2024-10-25 · Xiangyu Zeng, Kunchang Li, Chenting Wang, Xinhao Li 외

Multimodal Large Language Models (MLLMs) have demonstrated impressive performance in short video understanding. However, understanding long-form videos still remains challenging for MLLMs. This paper proposes TimeSuite, …

EgoSchemaHallucinationHighlight DetectionMoment Retrieval+3

VideoChat-M1: Collaborative Policy Planning for Video Understanding via Multi-Agent Reinforcement Learning

2025-11-24 · Boyu Chen, Zikang Wang, Zhengrong Yue, Kainan Yan 외 arxiv

By leveraging tool-augmented Multimodal Large Language Models (MLLMs), multi-agent frameworks are driving progress in video understanding. However, most of them adopt static and non-learnable tool invocation mechanisms, …

Multi-agent Reinforcement Learning

Online Video Understanding: OVBench and VideoChat-Online

2024-12-31 · CVPR 2025 1 · Zhenpeng Huang, Xinhao Li, Jiaqi Li, Jing Wang 외

Multimodal Large Language Models (MLLMs) have significantly progressed in offline video understanding. However, applying these models to real-world scenarios, such as autonomous driving and human-computer interaction, pr…

Autonomous DrivingQuestion AnsweringVideo Question AnsweringVideo Understanding

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

2025-04-09 · Xinhao Li, Ziang Yan, Desen Meng, Lu Dong 외

Recent advancements in reinforcement learning have significantly advanced the reasoning capabilities of multimodal large language models (MLLMs). While approaches such as Group Relative Policy Optimization (GRPO) and rul…

MVBenchObject TrackingVideo Understanding