paper-with-me

홈 › Papers

Leveraging Vision-Language Large Models for Interpretable Video Action Recognition with Semantic Tokenization

2025-09-06 · Jingwei Peng, Zhixuan Qiu, Boyu Jin, Surasakdi Siripong arxiv

Human action recognition often struggles with deep semantic understanding, complex contextual information, and fine-grained distinction, limitations that traditional methods frequently encounter when dealing with diverse video data. Inspired by the remarkable capabilities of large language models, this paper introduces LVLM-VAR, a novel framework that pioneers the application of pre-trained Vision-Language Large Models (LVLMs) to video action recognition, emphasizing enhanced accuracy and interpretability. Our method features a Video-to-Semantic-Tokens (VST) Module, which innovatively transforms raw video sequences into discrete, semantically and temporally consistent "semantic action tokens," effectively crafting an "action narrative" that is comprehensible to an LVLM. These tokens, combined with natural language instructions, are then processed by a LoRA-fine-tuned LVLM (e.g., LLaVA-13B) for robust action classification and semantic reasoning. LVLM-VAR not only achieves state-of-the-art or highly competitive performance on challenging benchmarks such as NTU RGB+D and NTU RGB+D 120, demonstrating significant improvements (e.g., 94.1% on NTU RGB+D X-Sub and 90.0% on NTU RGB+D 120 X-Set), but also substantially boosts model interpretability by generating natural language explanations for its predictions.

📄 PDF Abstract BibTeX arXiv:2509.05695

Code (0)

등록된 구현이 없습니다.

Tasks

Action ClassificationAction Recognition

Similar Papers 제목 키워드 기반

Programmatic Video Prediction Using Large Language Models

2025-05-20 · Hao Tang, Kevin Ellis, Suhas Lohit, Michael J. Jones 외

The task of estimating the world model describing the dynamics of a real world process assumes immense importance for anticipating and preparing for future outcomes. For applications such as video surveillance, robotics …

Autonomous DrivingPredictionVideo GenerationVideo Prediction

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos

2025-08-21 · Kaining Li, Shuwei He, Zihan Xu arxiv

Human action recognition in long-term videos, characterized by complex backgrounds and subtle action differences, poses significant challenges for traditional deep learning models due to computational overhead, difficult…

Action ClassificationAction UnderstandingAction Recognition

SOVABench: A Vehicle Surveillance Action Retrieval Benchmark for Multimodal Large Language Models

2026-01-08 · Oriol Rabasseda, Zenjie Li, Kamal Nasrollahi, Sergio Escalera arxiv

Automatic identification of events and recurrent behavior analysis are critical for video surveillance. However, most existing content-based video retrieval benchmarks focus on scene-level similarity and do not evaluate …

Visual ReasoningVideo Retrieval

Clue Matters: Leveraging Latent Visual Clues to Empower Video Reasoning

2026-03-16 · Kaixin zhang, Xiaohe Li, Jiahao Li, Haohua Wu 외 arxiv

Multi-modal Large Language Models (MLLMs) have significantly advanced video reasoning, yet Video Question Answering (VideoQA) remains challenging due to its demand for temporal causal reasoning and evidence-grounded answ…

Video Question AnsweringAnswer Generation

Neuro-Symbolic Representations for Video Captioning: A Case for Leveraging Inductive Biases for Vision and Language

2020-11-18 · Hassan Akbari, Hamid Palangi, Jianwei Yang, Sudha Rao 외

Neuro-symbolic representations have proved effective in learning structure information in vision and language. In this paper, we propose a new model architecture for learning multi-modal neuro-symbolic representations fo…

Dictionary LearningDisentanglementVideo Captioning