paper-with-me

홈 › Papers

InstrAct: Towards Action-Centric Understanding in Instructional Videos

2026-04-09 · Zhuoyi Yang, Jiapeng Yu, Reuben Tan, Boyang Li, Huijuan Xu arxiv

Understanding instructional videos requires recognizing fine-grained actions and modeling their temporal relations, which remains challenging for current Video Foundation Models (VFMs). This difficulty stems from noisy web supervision and a pervasive "static bias", where models rely on objects rather than motion cues. To address this, we propose InstrAction, a pretraining framework for instructional videos' action-centric representations. We first introduce a data-driven strategy, which filters noisy captions and generates action-centric hard negatives to disentangle actions from objects during contrastive learning. At the visual feature level, an Action Perceiver extracts motion-relevant tokens from redundant video encodings. Beyond contrastive learning, we introduce two auxiliary objectives: Dynamic Time Warping alignment (DTW-Align) for modeling sequential temporal structure, and Masked Action Modeling (MAM) for strengthening cross-modal grounding. Finally, we introduce the InstrAct Bench to evaluate action-centric understanding, where our method consistently outperforms state-of-the-art VFMs on semantic reasoning, procedural logic, and fine-grained retrieval tasks.

📄 PDF Abstract BibTeX arXiv:2604.08762

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive Learning

Similar Papers 제목 키워드 기반

ProcObject-10K: Benchmarking Object-Centric Procedural Understanding in Instructional Videos

2025-12-03 · Wenliang Guo, Yu Kong arxiv

Procedural activities are fundamentally driven by object state transitions, yet existing instructional video benchmarks remain action-centric and cannot evaluate whether models reason about how objects evolve toward task…

Retrieval-Augmented Egocentric Video Captioning

2024-01-01 · CVPR 2024 1 · Jilan Xu, Yifei HUANG, Junlin Hou, Guo Chen 외

Understanding human actions from videos of first-person view poses significant challenges. Most prior approaches explore representation learning on egocentric videos only, while overlooking the potential benefit of explo…

Representation LearningRetrievalVideo Captioning

How to Make a BLT Sandwich? Learning to Reason towards Understanding Web Instructional Videos

2018-12-02 · Shaojie Wang, Wentian Zhao, Ziyi Kou, Chenliang Xu

Understanding web instructional videos is an essential branch of video understanding in two aspects. First, most existing video methods focus on short-term actions for a-few-second-long video clips; these methods are not…

Logical ReasoningQuestion AnsweringVideo Understanding

Bridge-Prompt: Towards Ordinal Action Understanding in Instructional Videos

2022-03-26 · CVPR 2022 1 · Muheng Li, Lei Chen, Yueqi Duan, Zhilan Hu 외

Action recognition models have shown a promising capability to classify human actions in short video clips. In a real scenario, multiple correlated human actions commonly occur in particular orders, forming semantically …

Action SegmentationAction UnderstandingActivity RecognitionHuman Activity Recognition

AssistQ: Affordance-centric Question-driven Task Completion for Egocentric Assistant

2022-03-08 · Benita Wong, Joya Chen, You Wu, Stan Weixian Lei 외

A long-standing goal of intelligent assistants such as AR glasses/robots has been to assist users in affordance-centric real-world scenarios, such as "how can I run the microwave for 1 minute?". However, there is still n…

Visual Question Answering (VQA)