paper-with-me

Papers

iMOVE: Instance-Motion-Aware Video Understanding

2025-02-17 · Jiaze Li, Yaya Shi, Zongyang Ma, Haoran Xu, Feng Cheng, Huihui Xiao, Ruiwen Kang, Fan Yang, Tingting Gao, Di Zhang

Enhancing the fine-grained instance spatiotemporal motion perception capabilities of Video Large Language Models is crucial for improving their temporal and general video understanding. However, current models struggle to perceive detailed and complex instance motions. To address these challenges, we have made improvements from both data and model perspectives. In terms of data, we have meticulously curated iMOVE-IT, the first large-scale instance-motion-aware video instruction-tuning dataset. This dataset is enriched with comprehensive instance motion annotations and spatiotemporal mutual-supervision tasks, providing extensive training for the model's instance-motion-awareness. Building on this foundation, we introduce iMOVE, an instance-motion-aware video foundation model that utilizes Event-aware Spatiotemporal Efficient Modeling to retain informative instance spatiotemporal motion details while maintaining computational efficiency. It also incorporates Relative Spatiotemporal Position Tokens to ensure awareness of instance spatiotemporal positions. Evaluations indicate that iMOVE excels not only in video temporal understanding and general video understanding but also demonstrates significant advantages in long-term video understanding.

📄 PDF Abstract BibTeX arXiv:2502.11594

Code (0)

등록된 구현이 없습니다.

Tasks

Computational EfficiencyVideo Understanding

Similar Papers 제목 키워드 기반

UniMove: A Unified Model for Multi-city Human Mobility Prediction

2025-08-09 · Chonghua Han, Yuan Yuan, Yukun Liu, Jingtao Ding 외 arxiv

Human mobility prediction is vital for urban planning, transportation optimization, and personalized services. However, the inherent randomness, non-uniform time intervals, and complex patterns of human mobility, compoun…

InterRVOS: Interaction-aware Referring Video Object Segmentation

2025-06-03 · Woojeong Jin, Seongchan Kim, Seungryong Kim

Referring video object segmentation aims to segment the object in a video corresponding to a given natural language expression. While prior works have explored various referring scenarios, including motion-centric or mul…

8kObjectReferring Video Object SegmentationSemantic Segmentation+3

The TYC Dataset for Understanding Instance-Level Semantics and Motions of Cells in Microstructures

2023-08-23 · Christoph Reich, Tim Prangemeier, Heinz Koeppl

Segmenting cells and tracking their motion over time is a common task in biomedical applications. However, predicting accurate instance-wise segmentation and cell motions from microscopy imagery remains a challenging tas…

Efficient Motion-Aware Video MLLM

2025-01-01 · CVPR 2025 1 · Zijia Zhao, Yuqi Huo, Tongtian Yue, Longteng Guo 외

Most current video MLLMs rely on uniform frame sampling and image-level encoders, resulting in inefficient data processing and limited motion awareness. To address these challenges, we introduce EMA, an Efficient Mot…

Question AnsweringVideo Question AnsweringVideo Understanding

3D-Aware Instance Segmentation and Tracking in Egocentric Videos

2024-08-19 · Yash Bhalgat, Vadim Tschernezki, Iro Laina, João F. Henriques 외

Egocentric videos present unique challenges for 3D scene understanding due to rapid camera motion, frequent object occlusions, and limited object visibility. This paper introduces a novel approach to instance segmentatio…

3D Object ReconstructionInstance SegmentationObjectObject Reconstruction+5