paper-with-me

홈 › Papers

Where Do We (Not) Need Temporal Context in Low-Resource Video Task Adaptation?

2026-06-02 · Luc P. J. Sträter, Hazel Doughty arxiv

Parameter-efficient fine-tuning (PEFT) and probing enable adaptation of foundation models using only a small number of trainable parameters, making it attractive for video understanding where annotation and computation are expensive. However, video PEFT has focused on adapting image-pretrained models, while standard PEFT methods can also be applied to video representations. These settings are rarely compared and both confine temporal reasoning to a single component of the model, leaving open how temporal context should be distributed across backbone, PEFT and probe. In this work we provide a systematic study of model adaptation strategies for video understanding. We evaluate methods across appearance-focused, motion-focused and spatially dense settings, with a particular focus on scenarios with limited data where parameter-efficiency is most beneficial. Our results provide new insights into PEFT and probing across settings and demonstrate the importance of temporal context allocation for effective video adaptation

📄 PDF Abstract BibTeX arXiv:2606.03837

Code (0)

등록된 구현이 없습니다.

Tasks

parameter-efficient fine-tuning

Similar Papers 제목 키워드 기반

How Much Temporal Long-Term Context is Needed for Action Segmentation?

2023-08-22 · ICCV 2023 1 · Emad Bahrami, Gianpiero Francesca, Juergen Gall

Modeling long-term context in videos is crucial for many fine-grained tasks including temporal action segmentation. An interesting question that is still open is how much long-term temporal context is needed for optimal …

Action SegmentationSegmentationTemporal Action Segmentation

LongEgoRefer: A Benchmark for Long-Form Egocentric Video Referring Expression Comprehension

2026-07-02 · Shunya Kato, Taiki Miyanishi, Shuhei Kurita, Mahiro Ukai 외 arxiv

Egocentric videos capture rich and diverse human-object interactions and have emerged as a fundamental resource for understanding human activities related to objects. In this context, Video Referring Expression Comprehen…

Referring Expression

MMTF: Multi-Modal Temporal Fusion for Commonsense Video Question Answering

2023-10-06 · Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) Workshops, 2023 2023 10 · Mobeen Ahmad, Geonwoo Park, Dongchan Park, Sanguk Park

Video question answering is a challenging task that requires understanding the video and question in the same context. This becomes even harder when the questions involve reasoning, such as predicting future events or ex…

counterfactualQuestion AnsweringVideo Question Answering

Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence

2025-10-23 · Jiahao Meng, Xiangtai Li, Haochen Wang, Yue Tan 외 arxiv

Most video reasoning models only generate textual reasoning traces without indicating when and where key evidence appears. Recent models such as OpenAI-o3 have sparked wide interest in evidence-centered reasoning for ima…

Reinforcement Learning

To Find Where You Talk: Temporal Sentence Localization in Video with Attention Based Location Regression

2018-04-19 · Yitian Yuan, Tao Mei, Wenwu Zhu

Given an untrimmed video and a sentence description, temporal sentence localization aims to automatically determine the start and end points of the described sentence within the video. The problem is challenging as it ne…

regressionSentenceTemporal Localization