paper-with-me

Papers

PcmNet: Position-Sensitive Context Modeling Network for Temporal Action Localization

2021-03-09 · Xin Qin, Hanbin Zhao, Guangchen Lin, Hao Zeng, Songcen Xu, Xi Li

Temporal action localization is an important and challenging task that aims to locate temporal regions in real-world untrimmed videos where actions occur and recognize their classes. It is widely acknowledged that video context is a critical cue for video understanding, and exploiting the context has become an important strategy to boost localization performance. However, previous state-of-the-art methods focus more on exploring semantic context which captures the feature similarity among frames or proposals, and neglect positional context which is vital for temporal localization. In this paper, we propose a temporal-position-sensitive context modeling approach to incorporate both positional and semantic information for more precise action localization. Specifically, we first augment feature representations with directed temporal positional encoding, and then conduct attention-based information propagation, in both frame-level and proposal-level. Consequently, the generated feature representations are significantly empowered with the discriminative capability of encoding the position-aware context information, and thus benefit boundary detection and proposal evaluation. We achieve state-of-the-art performance on both two challenging datasets, THUMOS-14 and ActivityNet-1.3, demonstrating the effectiveness and generalization ability of our method.

📄 PDF Abstract BibTeX arXiv:2103.05270

Code (0)

등록된 구현이 없습니다.

Tasks

Action LocalizationBoundary DetectionPositionTemporal Action LocalizationTemporal LocalizationVideo Understanding

Similar Papers 제목 키워드 기반

Beyond Patches: Mining Interpretable Part-Prototypes for Explainable AI

2025-04-16 · Mahdi Alehdaghi, Rajarshi Bhattacharya, Pourya Shamsolmoali, Rafael M. O. Cruz 외

Deep learning has provided considerable advancements for multimedia systems, yet the interpretability of deep models remains a challenge. State-of-the-art post-hoc explainability methods, such as GradCAM, provide visual …

Unsupervised Part Discovery

Context Matters: An Empirical Study of the Impact of Contextual Information in Temporal Question Answering Systems

2024-06-27 · Dan Schumacher, Fatemeh Haji, Tara Grey, Niharika Bandlamudi 외

Large language models (LLMs) often struggle with temporal reasoning, crucial for tasks like historical event analysis and time-sensitive information retrieval. Despite advancements, state-of-the-art models falter in hand…

Information RetrievalQuestion AnsweringRetrieval

EgoAction: Egocentric Action Composition with Reliability-Aware Temporal Fusion for the EPIC-KITCHENS Action Detection Challenge at CVPR 2026

2026-05-23 · Zhiheng Fu, Zixu Li, Zhiwei Chen, Fangxu Liu 외 arxiv

The EPIC-KITCHENS-100 Action Detection challenge evaluates whether a model can localize the start and end of each action in long untrimmed egocentric videos and assign the corresponding verb--noun action label. In this r…

Action Detection

Unlearning Is Not Just Erasing: Temporal Decoupling via Generation Inequality

2026-08-24 · Xunlei Chen, Qirui Ye, Yuang Li, Yi Gong 외 arxiv

Large language models (LLMs) require effective unlearning to address privacy regulations and safety concerns. However, achieving precise forgetting without compromising general utility remains challenging. Existing seque…

Crowdsourcing a Large Dataset of Domain-Specific Context-Sensitive Semantic Verb Relations

2016-05-01 · LREC 2016 5 · Maria Sukhareva, Judith Eckle-Kohler, Ivan Habernal, Iryna Gurevych

We present a new large dataset of 12403 context-sensitive verb relations manually annotated via crowdsourcing. These relations capture fine-grained semantic information between verb-centric propositions, such as temporal…