paper-with-me

Papers Video-based Generative Performance Benchmarking

“Video-based Generative Performance Benchmarking” 태그가 달린 논문 20편 · 필터 해제

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models

2024-11-17 · Tingyu Qu, Mingxiao Li, Tinne Tuytelaars, Marie-Francine Moens

Recent advances in multimodal Large Language Models (LLMs) have shown great success in understanding multi-modal contents. For video understanding tasks, training-based video LLMs are difficult to build due to the scarci…

MVBenchVideo-based Generative Performance BenchmarkingVideo-based Generative Performance Benchmarking (Consistency)Video-based Generative Performance Benchmarking (Contextual Understanding)+5

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance

2024-11-04 · Ruyang Liu, Haoran Tang, Haibo Liu, Yixiao Ge 외

The past year has witnessed the significant advancement of video-based large language models. However, the challenge of developing a unified model for both short and long video understanding remains unresolved. Most exis…

Caption GenerationMultiple-choiceVideo-based Generative Performance BenchmarkingVideo-based Generative Performance Benchmarking (Consistency)+7

SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

2024-07-22 · Mingze Xu, Mingfei Gao, Zhe Gan, Hong-You Chen 외

We propose SlowFast-LLaVA (or SF-LLaVA for short), a training-free video large language model (LLM) that can jointly capture detailed spatial semantics and long-range temporal context without exceeding the token budget o…

Language ModelingLanguage ModellingLarge Language Model+8

VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding

2024-06-13 · Muhammad Maaz, Hanoona Rasheed, Salman Khan, Fahad Khan

Building on the advances of language models, Large Multimodal Models (LMMs) have contributed significant improvements in video understanding. While the current video LMMs utilize advanced Large Language Models (LLMs), th…

Dense Video CaptioningMVBenchQuestion AnsweringVCGBench-Diverse+10

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

2024-04-25 · arXiv 2024 4 · Lin Xu, Yilin Zhao, Daquan Zhou, Zhijie Lin 외

Vision-language pre-training has significantly elevated performance across a wide range of image-language applications. Yet, the pre-training process for video-related tasks demands exceptionally large computational and …

Dense CaptioningMVBenchVideo-based Generative Performance BenchmarkingVideo-based Generative Performance Benchmarking (Consistency)+7

ST-LLM: Large Language Models Are Effective Temporal Learners

2024-03-30 · Ruyang Liu, Chen Li, Haoran Tang, Yixiao Ge 외

Large Language Models (LLMs) have showcased impressive capabilities in text comprehension and generation, prompting research efforts towards video LLMs to facilitate human-AI interaction at the video level. However, how …

MVBenchReading ComprehensionVideo-based Generative Performance BenchmarkingVideo Question Answering+1

An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM

2024-03-27 · Wonkyun Kim, Changin Choi, Wonseok Lee, Wonjong Rhee

Stimulated by the sophisticated reasoning capabilities of recent Large Language Models (LLMs), a variety of strategies for bridging video modality have been devised. A prominent strategy involves Video Language Models (V…

Language ModelingLanguage ModellingMultiple-choiceQuestion Answering+3

LITA: Language Instructed Temporal-Localization Assistant

2024-03-27 · De-An Huang, Shijia Liao, Subhashree Radhakrishnan, Hongxu Yin 외

There has been tremendous progress in multimodal Large Language Models (LLMs). Recent works have extended these models to video input with promising instruction following capabilities. However, an important missing piece…

Instruction FollowingTemporal LocalizationText GenerationVideo-based Generative Performance Benchmarking+1

CAT: Enhancing Multimodal Large Language Model to Answer Questions in Dynamic Audio-Visual Scenarios

2024-03-07 · Qilang Ye, Zitong Yu, Rui Shao, Xinyu Xie 외

This paper focuses on the challenge of answering questions in scenarios that are composed of rich and complex dynamic audio-visual components. Although existing Multimodal Large Language Models (MLLMs) can respond to aud…

Audio-visual Question AnsweringAudio-Visual Question Answering (AVQA)Language ModelingLanguage Modelling+6

Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback

2024-02-06 · Daechul Ahn, Yura Choi, Youngjae Yu, Dongyeop Kang 외

Recent advancements in large language models have influenced the development of video large multimodal models (VLMMs). The previous approaches for VLMMs involved Supervised Fine-Tuning (SFT) with instruction-tuned datase…

Video-based Generative Performance Benchmarking

VTimeLLM: Empower LLM to Grasp Video Moments

2023-11-30 · CVPR 2024 1 · Bin Huang, Xin Wang, Hong Chen, Zihan Song 외

Large language models (LLMs) have shown remarkable text understanding capabilities, which have been extended as Video LLMs to handle video data for comprehending visual details. However, existing Video LLMs can only prov…

Dense Video CaptioningTemporal Relation ExtractionVCGBench-DiverseVideo-based Generative Performance Benchmarking+8

LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

2023-11-28 · Yanwei Li, Chengyao Wang, Jiaya Jia

In this work, we present a novel method to tackle the token generation challenge in Vision Language Models (VLMs) for video and image understanding, called LLaMA-VID. Current VLMs, while proficient in tasks like image ca…

Image CaptioningQuestion AnsweringVideo-based Generative Performance BenchmarkingVideo Question Answering+2

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

2023-11-28 · CVPR 2024 1 · Kunchang Li, Yali Wang, Yinan He, Yizhuo Li 외

With the rapid development of Multi-modal Large Language Models (MLLMs), a number of diagnostic benchmarks have recently emerged to evaluate the comprehension capabilities of these models. However, most benchmarks predom…

3D Question Answering (3D-QA)DiagnosticFairnessMultiple-choice+12

Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding

2023-11-14 · CVPR 2024 1 · Peng Jin, Ryuichi Takanobu, Wancai Zhang, Xiaochun Cao 외

Large language models have demonstrated impressive universal capabilities across a wide range of open-ended tasks and have extended their utility to encompass multimodal conversations. However, existing methods encounter…

Image-based Generative Performance BenchmarkingLanguage ModelingLanguage ModellingScience Question Answering+11

BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning

2023-09-27 · CVPR 2024 1 · Ruyang Liu, Chen Li, Yixiao Ge, Ying Shan 외

The recent progress in Large Language Models (LLM) has spurred various advancements in image-language conversation agents, while how to build a proficient video-based dialogue system is still under exploration. Consideri…

GPUVideo-based Generative Performance BenchmarkingVideo-based Generative Performance Benchmarking (Consistency)Video-based Generative Performance Benchmarking (Contextual Understanding)+7

MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

2023-07-31 · CVPR 2024 1 · Enxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang 외

Recently, integrating video foundation models and large language models to build a video understanding system can overcome the limitations of specific pre-defined vision tasks. Yet, existing systems can only handle video…

Multiple-choiceQuestion AnsweringVideo-based Generative Performance BenchmarkingVideo-based Generative Performance Benchmarking (Consistency)+12

Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

2023-06-08 · Muhammad Maaz, Hanoona Rasheed, Salman Khan, Fahad Shahbaz Khan

Conversation agents fueled by Large Language Models (LLMs) are providing a new way to interact with visual data. While there have been initial attempts for image-based conversation models, this work addresses the under-e…

Question AnsweringVCGBench-DiverseVideo-based Generative Performance BenchmarkingVideo-based Generative Performance Benchmarking (Consistency)+7

Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

2023-06-05 · Hang Zhang, Xin Li, Lidong Bing

We present Video-LLaMA a multi-modal framework that empowers Large Language Models (LLMs) with the capability of understanding both visual and auditory content in the video. Video-LLaMA bootstraps cross-modal training fr…

Language ModelingLanguage ModellingText GenerationVideo-based Generative Performance Benchmarking+9

VideoChat: Chat-Centric Video Understanding

2023-05-10 · Kunchang Li, Yinan He, Yi Wang, Yizhuo Li 외

In this paper, we initiate an attempt of developing an end-to-end chat-centric video understanding system, coined as VideoChat. It integrates video foundation models and large language models via a learnable neural inter…

Question AnsweringVideo-based Generative Performance BenchmarkingVideo-based Generative Performance Benchmarking (Consistency)Video-based Generative Performance Benchmarking (Contextual Understanding)+6

LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

2023-04-28 · Peng Gao, Jiaming Han, Renrui Zhang, Ziyi Lin 외

How to efficiently transform large language models (LLMs) into instruction followers is recently a popular research direction, while training LLM for multi-modal reasoning remains less explored. Although the recent LLaMA…

Instruction FollowingmodelOptical Character Recognition (OCR)Video-based Generative Performance Benchmarking+9
1–20 / 20