paper-with-me

홈 › Papers

SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding

2025-03-24 · Mingze Xu, Mingfei Gao, Shiyu Li, Jiasen Lu, Zhe Gan, Zhengfeng Lai, Meng Cao, Kai Kang, Yinfei Yang, Afshin Dehghan

We introduce SlowFast-LLaVA-1.5 (abbreviated as SF-LLaVA-1.5), a family of video large language models (LLMs) offering a token-efficient solution for long-form video understanding. We incorporate the two-stream SlowFast mechanism into a streamlined training pipeline, and perform joint video-image training on a carefully curated data mixture of only publicly available datasets. Our primary focus is on highly efficient model scales (1B and 3B), demonstrating that even relatively small Video LLMs can achieve state-of-the-art performance on video understanding, meeting the demand for mobile-friendly models. Experimental results demonstrate that SF-LLaVA-1.5 achieves superior performance on a wide range of video and image tasks, with robust results at all model sizes (ranging from 1B to 7B). Notably, SF-LLaVA-1.5 achieves state-of-the-art results in long-form video understanding (e.g., LongVideoBench and MLVU) and excels at small scales across various video benchmarks.

📄 PDF Abstract BibTeX arXiv:2503.18943

Code (0)

등록된 구현이 없습니다.

Tasks

FormVideo Understanding

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

2024-07-22 · Mingze Xu, Mingfei Gao, Zhe Gan, Hong-You Chen 외

We propose SlowFast-LLaVA (or SF-LLaVA for short), a training-free video large language model (LLM) that can jointly capture detailed spatial semantics and long-range temporal context without exceeding the token budget o…

Language ModelingLanguage ModellingLarge Language Model+8

LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

2025-01-07 · Shaolei Zhang, Qingkai Fang, Zhe Yang, Yang Feng

The advent of real-time large multimodal models (LMMs) like GPT-4o has sparked considerable interest in efficient LMMs. LMM frameworks typically encode visual inputs into vision tokens (continuous representations) and in…

GPUVisual Question Answering (VQA)Zero-Shot Video Question Answer

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs

2025-06-27 · Boyuan Sun, Jiaxing Zhao, Xihan Wei, Qibin Hou

In this paper, we present LLaVA-Scissor, a training-free token compression strategy designed for video multimodal large language models. Previous methods mostly attempt to compress tokens based on attention scores, but f…

Question AnsweringVideo Question AnsweringVideo Understanding

LLaVA-OneVision: Easy Visual Task Transfer

2024-08-06 · Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang 외

We present LLaVA-OneVision, a family of open large multimodal models (LMMs) developed by consolidating our insights into data, models, and visual representations in the LLaVA-NeXT blog series. Our experimental results de…

3D Question Answering (3D-QA)Multiple-choiceTemporal Relation Extraction+6

ForestPrune: High-ratio Visual Token Compression for Video Multimodal Large Language Models via Spatial-Temporal Forest Modeling

2026-03-24 · Shaobo Ju, Baiyang Song, Tao Chen, Jiapeng Zhang 외 arxiv

Due to the great saving of computation and memory overhead, token compression has become a research hot-spot for MLLMs and achieved remarkable progress in image-language tasks. However, for the video, existing methods st…