paper-with-me

홈 › Papers

Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives

2024-06-09 · Thong Nguyen, Yi Bin, Junbin Xiao, Leigang Qu, Yicong Li, Jay Zhangjie Wu, Cong-Duy Nguyen, See-Kiong Ng, Luu Anh Tuan

Humans use multiple senses to comprehend the environment. Vision and language are two of the most vital senses since they allow us to easily communicate our thoughts and perceive the world around us. There has been a lot of interest in creating video-language understanding systems with human-like senses since a video-language pair can mimic both our linguistic medium and visual environment with temporal dynamics. In this survey, we review the key tasks of these systems and highlight the associated challenges. Based on the challenges, we summarize their methods from model architecture, model training, and data perspectives. We also conduct performance comparison among the methods, and discuss promising directions for future research.

📄 PDF Abstract BibTeX arXiv:2406.05615

Code (1)

nguyentthong/video-language-understanding 공식 구현 tf

Tasks

model

Similar Papers 제목 키워드 기반

Large Language Models for Crash Detection in Video: A Survey of Methods, Datasets, and Challenges

2025-07-02 · Sanjeda Akter, Ibne Farabi Shihab, Anuj Sharma

Crash detection from video feeds is a critical problem in intelligent transportation systems. Recent developments in large language models (LLMs) and vision-language models (VLMs) have transformed how we process, reason …

Video Understanding

VideoLLM Benchmarks and Evaluation: A Survey

2025-05-03 · Yogesh Kumar

The rapid development of Large Language Models (LLMs) has catalyzed significant advancements in video understanding technologies. This survey provides a comprehensive analysis of benchmarks and evaluation methodologies s…

SurveyVideo Understanding

Video Understanding by Design: How Datasets Shape Video Models

2025-09-11 · Lei Wang, Syuan-Hao Li, Piotr Koniusz, Yongsheng Gao arxiv

Research in video understanding has advanced rapidly, driven by increasingly diverse datasets and more powerful model architectures. While existing surveys typically organize progress by tasks, benchmarks, or model famil…

Image-to-Video Transfer Learning based on Image-Language Foundation Models: A Comprehensive Survey

2025-10-12 · Jinxuan Li, Chaolei Tan, Haoxuan Chen, Jianxin Ma 외 arxiv

Image-Language Foundation Models (ILFMs) have demonstrated remarkable success in vision-language understanding, providing transferable multimodal representations that generalize across diverse downstream image-based task…

Spatio-Temporal Video GroundingVideo Question AnsweringTransfer Learning

Video Understanding with Large Language Models: A Survey

2023-12-29 · Yunlong Tang, Jing Bi, Siting Xu, Luchuan Song 외

With the burgeoning growth of online video platforms and the escalating volume of video content, the demand for proficient video understanding tools has intensified markedly. Given the remarkable capabilities of large la…

SurveyVideo Understanding