paper-with-me

홈 › Papers

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos

2025-08-21 · Kaining Li, Shuwei He, Zihan Xu arxiv

Human action recognition in long-term videos, characterized by complex backgrounds and subtle action differences, poses significant challenges for traditional deep learning models due to computational overhead, difficulty in capturing long-range temporal dependencies, and limited semantic understanding. While Large Language Models (LLMs) and Large Vision-Language Models (LVLMs) have shown remarkable capabilities in multi-modal understanding and reasoning, their direct application to continuous video streams for fine-grained action recognition remains an open problem. This paper introduces VT-LVLM-AR (Video-Temporal Large Vision-Language Model Adapter for Action Recognition), a novel framework designed to bridge this gap. VT-LVLM-AR comprises a Video-to-Event Mapper (VTEM) that efficiently transforms raw video into compact, semantically rich, and temporally coherent "visual event sequences" through lightweight spatio-temporal feature extraction, adaptive temporal pooling, and conceptual quantization with an event coherence bias. These visual event sequences are then fed into an LVLM-based Action Reasoning module, specifically a frozen LLaVA-1.5 model, adapted using parameter-efficient Prompt Tuning (P-Tuning v2) for action classification. Comprehensive evaluations on the NTU RGB+D and NTU RGB+D 120 datasets demonstrate that VT-LVLM-AR consistently achieves state-of-the-art performance, surpassing existing methods (e.g., 94.1% accuracy on NTU RGB+D X-Sub). Ablation studies confirm the critical contributions of VTEM's components and the efficacy of Prompt Tuning, while human evaluations underscore the interpretability of our visual event representations. This work highlights the immense potential of leveraging LVLMs for robust and interpretable video action understanding through effective video-to-language translation and efficient model adaptation.

📄 PDF Abstract BibTeX arXiv:2508.15903

Code (0)

등록된 구현이 없습니다.

Tasks

Action ClassificationAction UnderstandingAction Recognition

Results from the Paper

RankTaskDatasetModelMetrics
#12 Action Recognition NTU RGB+D VT-LVLM-AR Accuracy (CS): 94.1

Similar Papers 제목 키워드 기반

Temporal-Oriented Recipe for Transferring Large Vision-Language Model to Video Understanding

2025-05-19 · Thong Nguyen, Zhiyuan Hu, Xu Lin, Cong-Duy Nguyen 외

Recent years have witnessed outstanding advances of large vision-language models (LVLMs). In order to tackle video understanding, most of them depend upon their implicit temporal understanding capacity. As such, they hav…

Language ModelingLanguage ModellingLarge Language ModelVideo Understanding

Fewer Tokens and Fewer Videos: Extending Video Understanding Abilities in Large Vision-Language Models

2024-06-12 · Shimin Chen, Yitian Yuan, Shaoxiang Chen, Zequn Jie 외

Amidst the advancements in image-based Large Vision-Language Models (image-LVLM), the transition to video-based models (video-LVLM) is hindered by the limited availability of quality video data. This paper addresses the …

Video Understanding

TempJail: Temporal Jailbreak Attack against Large Vision-Language Models via Subtitle Scheduling

2026-08-20 · Ling Zhou, Yihao Huang, Jingling Sun, Zhiwen Tian 외 arxiv

Large vision-language models (LVLMs) have achieved remarkable progress in video understanding and reasoning. Despite extensive studies on text- and image-based jailbreaks, video jailbreaks against LVLMs remain largely un…

GenVideoLens: Where LVLMs Fall Short in AI-Generated Video Detection?

2026-03-19 · Yueying Zou, Pei Pei Li, Zekun Li, Xinyu Guo 외 arxiv

In recent years, AI-generated videos have become increasingly realistic and sophisticated. Meanwhile, Large Vision-Language Models (LVLMs) have shown strong potential for detecting such content. However, existing evaluat…

Binary Classification

SVBench: A Benchmark with Temporal Multi-Turn Dialogues for Streaming Video Understanding

2025-02-15 · Zhenyu Yang, Yuhang Hu, Zemin Du, Dizhan Xue 외

Despite the significant advancements of Large Vision-Language Models (LVLMs) on established benchmarks, there remains a notable gap in suitable evaluation regarding their applicability in the emerging domain of long-cont…

Question AnsweringStreaming video understandingVideo Understanding