paper-with-me

Video-based Generative Performance Benchmarking (Temporal Understanding)

1개 벤치마크 · 논문 15편 · 이 태스크의 논문 보기 →

Benchmarks

VideoInstruct

결과 54개

Most implemented

Papers

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models

2024-11-17 · Tingyu Qu, Mingxiao Li, Tinne Tuytelaars, Marie-Francine Moens

Recent advances in multimodal Large Language Models (LLMs) have shown great success in understanding multi-modal contents. For video understanding tasks, training-based video LLMs are difficult to build due to the scarci…

MVBenchVideo-based Generative Performance BenchmarkingVideo-based Generative Performance Benchmarking (Consistency)Video-based Generative Performance Benchmarking (Contextual Understanding)+5

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance

2024-11-04 · Ruyang Liu, Haoran Tang, Haibo Liu, Yixiao Ge 외

The past year has witnessed the significant advancement of video-based large language models. However, the challenge of developing a unified model for both short and long video understanding remains unresolved. Most exis…

Caption GenerationMultiple-choiceVideo-based Generative Performance BenchmarkingVideo-based Generative Performance Benchmarking (Consistency)+7

SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

2024-07-22 · Mingze Xu, Mingfei Gao, Zhe Gan, Hong-You Chen 외

We propose SlowFast-LLaVA (or SF-LLaVA for short), a training-free video large language model (LLM) that can jointly capture detailed spatial semantics and long-range temporal context without exceeding the token budget o…

Language ModelingLanguage ModellingLarge Language Model+8

VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding

2024-06-13 · Muhammad Maaz, Hanoona Rasheed, Salman Khan, Fahad Khan

Building on the advances of language models, Large Multimodal Models (LMMs) have contributed significant improvements in video understanding. While the current video LMMs utilize advanced Large Language Models (LLMs), th…

Dense Video CaptioningMVBenchQuestion AnsweringVCGBench-Diverse+10

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

2024-04-25 · arXiv 2024 4 · Lin Xu, Yilin Zhao, Daquan Zhou, Zhijie Lin 외

Vision-language pre-training has significantly elevated performance across a wide range of image-language applications. Yet, the pre-training process for video-related tasks demands exceptionally large computational and …

Dense CaptioningMVBenchVideo-based Generative Performance BenchmarkingVideo-based Generative Performance Benchmarking (Consistency)+7

MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

2024-04-04 · Kirolos Ataallah, Xiaoqian Shen, Eslam Abdelrahman, Essam Sleiman 외

This paper introduces MiniGPT4-Video, a multimodal Large Language Model (LLM) designed specifically for video understanding. The model is capable of processing both temporal visual and textual data, making it adept at un…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+10

전체 15편 보기 →