paper-with-me

Papers

MALT: Multi-scale Action Learning Transformer for Online Action Detection

2024-05-31 · Zhipeng Yang, Ruoyu Wang, Yang Tan, Liping Xie

Online action detection (OAD) aims to identify ongoing actions from streaming video in real-time, without access to future frames. Since these actions manifest at varying scales of granularity, ranging from coarse to fine, projecting an entire set of action frames to a single latent encoding may result in a lack of local information, necessitating the acquisition of action features across multiple scales. In this paper, we propose a multi-scale action learning transformer (MALT), which includes a novel recurrent decoder (used for feature fusion) that includes fewer parameters and can be trained more efficiently. A hierarchical encoder with multiple encoding branches is further proposed to capture multi-scale action features. The output from the preceding branch is then incrementally input to the subsequent branch as part of a cross-attention calculation. In this way, output features transition from coarse to fine as the branches deepen. We also introduce an explicit frame scoring mechanism employing sparse attention, which filters irrelevant frames more efficiently, without requiring an additional network. The proposed method achieved state-of-the-art performance on two benchmark datasets (THUMOS'14 and TVSeries), outperforming all existing models used for comparison, with an mAP of 0.2% for THUMOS'14 and an mcAP of 0.1% for TVseries.

📄 PDF Abstract BibTeX arXiv:2405.20892

Code (0)

등록된 구현이 없습니다.

Tasks

Action DetectionDecoderOnline Action Detection

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

MALT Diffusion: Memory-Augmented Latent Transformers for Any-Length Video Generation

2025-02-18 · Sihyun Yu, Meera Hahn, Dan Kondratyuk, Jinwoo Shin 외

Diffusion models are successful for synthesizing high-quality videos but are limited to generating short clips (e.g., 2-10 seconds). Synthesizing sustained footage (e.g. over minutes) still remains an open research quest…

Text-to-Video GenerationVideo Generation

A Social Opinion Gold Standard for the Malta Government Budget 2018

2019-11-01 · WS 2019 11 · Keith Cortis, Brian Davis

We present a gold standard of annotated social opinion for the Malta Government Budget 2018. It consists of over 500 online posts in English and/or the Maltese less-resourced language, gathered from social media platform…

NegationOpinion Mining

ThermalTap: Passive Application Fingerprinting in VR Headsets via Thermal Side Channels

2026-05-13 · Mahsin Bin Akram, A H M Nazmus Sakib, OFM Riaz Rahman Aranya, Raveen Wijewickrama 외 arxiv

Standalone virtual reality (VR) headsets process highly sensitive personal, professional, and health-related data, yet their susceptibility to non-contact physical side channels remains largely unexplored. Existing side-…

Malta National Language Technology Platform: A vision for enhancing Malta’s official languages using Machine Translation

2021-09-01 · MMTLRL (RANLP) 2021 9 · Keith Cortis, Judie Attard, Donatienne Spiteri

In this paper we introduce a vision towards establishing the Malta National Language Technology Platform; an ongoing effort that aims to provide a basis for enhancing Malta’s official languages, namely Maltese and Englis…

Machine TranslationTranslation

Crowd-sourcing evaluation of automatically acquired, morphologically related word groupings

2014-05-01 · LREC 2014 5 · Claudia Borg, Albert Gatt

The automatic discovery and clustering of morphologically related words is an important problem with several practical applications. This paper describes the evaluation of word clusters carried out through crowd-sourcing…

Clustering