paper-with-me

홈 › Papers

Temporal Gains, Spatial Costs: Revisiting Video Fine-Tuning in Multimodal Large Language Models

2026-03-18 · Linghao Zhang, Jungang Li, Yonghua Hei, Sicheng Tao, Song Dai, Yibo Yan, Zihao Dongfang, Weiting Liu, Chenxi Qin, Hanqian Li, Xin Zou, Jiahao Zhang, Shuhang Xun, Haiyun Jiang, Xuming Hu arxiv

Multimodal large language models (MLLMs) are typically trained in multiple stages, with video-based supervised fine-tuning (Video-SFT) serving as a key step for improving visual understanding. Yet its effect on the fine-grained evolution of visual capabilities, particularly the balance between spatial and temporal understanding, remains poorly understood. In this paper, we systematically study how Video-SFT reshapes visual capabilities in MLLMs. Across architectures, parameter scales, and frame sampling settings, we observe a consistent pattern: Video-SFT reliably improves video performance, but often yields limited gains or even degradation on static image benchmarks. We further show that this trade-off is closely tied to temporal budget: increasing the number of sampled frames generally improves video performance, but does not reliably improve static image performance. Motivated by this finding, we study an instruction-aware Hybrid-Frame strategy that adaptively allocates frame counts and partially mitigates the image-video trade-off. Our results indicate that Video-SFT is not a free lunch for MLLMs, and preserving spatial understanding remains a central challenge in joint image-video training.

📄 PDF Abstract BibTeX arXiv:2603.17541

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Revisiting Temporal Alignment for Video Restoration

2021-11-30 · CVPR 2022 1 · Kun Zhou, Wenbo Li, Liying Lu, Xiaoguang Han 외

Long-range temporal alignment is critical yet challenging for video restoration tasks. Recently, some works attempt to divide the long-range alignment into several sub-alignments and handle them progressively. Although t…

DeblurringDenoisingMotion CompensationSuper-Resolution+2

STOP: Integrated Spatial-Temporal Dynamic Prompting for Video Understanding

2025-03-20 · CVPR 2025 1 · Zichen Liu, Kunlun Xu, Bing Su, Xu Zou 외

Pre-trained on tremendous image-text pairs, vision-language models like CLIP have demonstrated promising zero-shot generalization across numerous image-based tasks. However, extending these capabilities to video tasks re…

Video UnderstandingZero-shot Generalization

FlowTrack: Revisiting Optical Flow for Long-Range Dense Tracking

2024-01-01 · CVPR 2024 1 · Seokju Cho, Jiahui Huang, Seungryong Kim, Joon-Young Lee

In the domain of video tracking existing methods often grapple with a trade-off between spatial density and temporal range. Current approaches in dense optical flow estimators excel in providing spatially dense track…

Optical Flow Estimation

Token Merging via Spatiotemporal Information Mining for Surgical Video Understanding

2025-09-28 · Xixi Jiang, Chen Yang, Dong Zhang, Pingcheng Dong 외 arxiv

Vision Transformer models have shown impressive effectiveness in the surgical video understanding tasks through long-range dependency modeling. However, current methods suffer from prohibitive computational costs due to …

Revisiting Temporal Modeling for CLIP-based Image-to-Video Knowledge Transferring

2023-01-26 · CVPR 2023 1 · Ruyang Liu, Jingjia Huang, Ge Li, Jiashi Feng 외

Image-text pretrained models, e.g., CLIP, have shown impressive general multi-modal knowledge learned from large-scale image-text data pairs, thus attracting increasing attention for their potential to improve visual rep…

Representation LearningRetrievalText RetrievalVideo Recognition+2