paper-with-me

홈 › Papers

Boundary-Centric Clip-Budgeted Active Learning for Temporal Action Segmentation

2026-04-16 · Halil Ismail Helvaci, Sen-ching Samson Cheung arxiv

Temporal action segmentation (TAS) in untrimmed videos requires dense temporal supervision. However, most of the annotation cost is spent identifying action transitions where segmentation errors concentrate and small temporal shifts can disproportionately degrade segment-level metrics. We introduce B-ACT, a clip-budgeted active learning framework that explicitly allocates supervision to these error-prone boundary regions. B-ACT operates in a hierarchical two-stage loop: (i) it ranks and queries unlabeled videos using predictive uncertainty, and (ii) within each selected video, it detects candidate transitions from the current model predictions and selects the top-$K$ boundaries via a novel boundary score. The boundary score fuses neighborhood uncertainty, class ambiguity, and temporal prediction dynamics to reveal the underlying importance of each frame. Importantly, our annotation protocol requests labels only at the boundary frames while still training on boundary-centered clips to exploit temporal context through the model's receptive field. Extensive experiments on GTEA, 50Salads, and Breakfast demonstrate that boundary-centric supervision delivers strong label efficiency and consistently surpasses representative TAS active learning baselines and prior state of the art under sparse budgets. Gains are largest on datasets where performance is highly sensitive to boundary placement, as measured by edit and overlap-based F1 metrics.

📄 PDF Abstract BibTeX arXiv:2604.15173

Code (0)

등록된 구현이 없습니다.

Tasks

Action SegmentationActive Learning

Similar Papers 제목 키워드 기반

ClipTBP: Clip-Pair based Temporal Boundary Prediction with Boundary-Aware Learning for Moment Retrieval

2026-04-30 · Ji-Hyeon Kim, Ho-Joong Kim, Seong-Whan Lee arxiv

Video moment retrieval is the task of retrieving specific segments of a video corresponding to a given text query. Recent studies have been conducted to improve multimodal alignment performance through visual-linguistic …

Moment Retrieval

Beyond Boundary Frames: Context-Centric Video Interpolation with Audio-Visual Semantics

2025-12-03 · Yuchen Deng, Xiuyang Wu, Hai-Tao Zheng, Jie Wang 외 arxiv

Video frame interpolation has long been challenged by limited controllability and interactivity, especially in scenarios involving fast, highly non-linear, and fine-grained motion. Although recent interactive interpolati…

Video Frame Interpolation

Enhancing Next Active Object-based Egocentric Action Anticipation with Guided Attention

2023-05-22 · Sanket Thakur, Cigdem Beyan, Pietro Morerio, Vittorio Murino 외

Short-term action anticipation (STA) in first-person videos is a challenging task that involves understanding the next active object interactions and predicting future actions. Existing action anticipation methods have p…

Action AnticipationObjectShort-term Object Interaction Anticipation

Guided Attention for Next Active Object @ EGO4D STA Challenge

2023-05-25 · Sanket Thakur, Cigdem Beyan, Pietro Morerio, Vittorio Murino 외

In this technical report, we describe the Guided-Attention mechanism based solution for the short-term anticipation (STA) challenge for the EGO4D challenge. It combines the object detections, and the spatiotemporal featu…

ObjectShort-term Object Interaction Anticipation

TIE: Time Interval Encoding for Video Generation over Events

2026-05-11 · Zhilei Shu, Shangwen Zhu, Zihang Liang, Xiaofan Li 외 arxiv

Director-style prompting, robotic action prediction, and interactive video agents demand temporal grounding over concurrent events -- a regime in which 68% of general clips and over 99% of robotics/gameplay clips contain…

Video Generation