paper-with-me

홈 › Papers

UFVideo: Towards Unified Fine-Grained Video Cooperative Understanding with Large Language Models

2025-12-12 · Hewen Pan, Cong Wei, Dashuang Liang, Zepeng Huang, Pengfei Gao, Ziqi Zhou, Lulu Xue, Pengfei Yan, Xiaoming Wei, Minghui Li, Shengshan Hu arxiv

With the advancement of multi-modal Large Language Models (LLMs), Video LLMs have been further developed to perform on holistic and specialized video understanding. However, existing works are limited to specialized video understanding tasks, failing to achieve a comprehensive and multi-grained video perception. To bridge this gap, we introduce UFVideo, the first Video LLM with unified multi-grained cooperative understanding capabilities. Specifically, we design unified visual-language guided alignment to flexibly handle video understanding across global, pixel and temporal scales within a single model. UFVideo dynamically encodes the visual and text inputs of different tasks and generates the textual response, temporal localization, or grounded mask. Additionally, to evaluate challenging multi-grained video understanding tasks, we construct the UFVideo-Bench consisting of three distinct collaborative tasks within the scales, which demonstrates UFVideo's flexibility and advantages over GPT-4o. Furthermore, we validate the effectiveness of our model across 9 public benchmarks covering various common video understanding tasks, providing valuable insights for future Video LLMs.

📄 PDF Abstract BibTeX arXiv:2512.11336

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Video-Text as Game Players: Hierarchical Banzhaf Interaction for Cross-Modal Representation Learning

2023-03-25 · CVPR 2023 1 · Peng Jin, Jinfa Huang, Pengfei Xiong, Shangxuan Tian 외

Contrastive learning-based video-language representation learning approaches, e.g., CLIP, have achieved outstanding performance, which pursue semantic interaction upon pre-defined video-text pairs. To clarify this coarse…

Contrastive LearningQuestion AnsweringRepresentation LearningRetrieval+3

WTS: A Pedestrian-Centric Traffic Video Dataset for Fine-grained Spatial-Temporal Understanding

2024-07-22 · Quan Kong, Yuki Kawana, Rajat Saini, Ashutosh Kumar 외

In this paper, we address the challenge of fine-grained video event understanding in traffic scenarios, vital for autonomous driving and safety. Traditional datasets focus on driver or vehicle behavior, often neglecting …

Autonomous Driving

Seeing Together: Multi-Robot Cooperative Egocentric Spatial Reasoning with Multimodal Large Language Models

2026-05-18 · Kunyu Peng, Zhikun Zhou, Kailun Yang, Di Wen 외 arxiv

Multimodal Large Language Models (MLLMs) have made substantial progress in egocentric video understanding, but their ability to reason cooperatively from multiple embodied viewpoints remains largely unexplored. We study …

Spatial Reasoning

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning

2024-12-30 · Peng Jin, Hao Li, Li Yuan, Shuicheng Yan 외

Multimodal representation learning, with contrastive learning, plays an important role in the artificial intelligence domain. As an important subfield, video-language representation learning focuses on learning represent…

Contrastive LearningQuestion AnsweringRepresentation LearningVideo Captioning+2

MotionAgent: Fine-grained Controllable Video Generation via Motion Field Agent

2025-02-05 · Xinyao Liao, Xianfang Zeng, Liao Wang, Gang Yu 외

We propose MotionAgent, enabling fine-grained motion control for text-guided image-to-video generation. The key technique is the motion field agent that converts motion information in text prompts into explicit motion fi…

Image to Video GenerationMotion GenerationOptical Flow EstimationVideo Generation