TrAction: Action Recognition with Sparse Trajectories
Modern action recognition models operate on memory- and compute-intensive dense RGB video volumes and frequently exploit appearance and background shortcuts, for example, predicting actions from objects or scenes instead of characteristic motion. We investigate an efficient alternative input modality that is largely free of such biases by construction: sparse point trajectories. To this end, we develop a simple transformer architecture for 2.5D trajectory-based recognition together with a masked-trajectory pretraining, which we show to substantially improve downstream action recognition accuracy. Despite using only a fraction of the dense RGB input, our method reaches 45% top-1 on Something-Something V2 and 54% on EPIC-Kitchens-100, and surpasses V-JEPA on time-reversal sensitivity. More importantly, we find trajectory features to be complementary to state-of-the-art appearance-based features. Fusing our pretrained model with DINOv2 and V-JEPA 2 improves top-1 accuracy on Something-Something V2 by 8.7 and 1.6 points, respectively. Code: https://github.com/ecker-lab/TrAction
Code (0)
등록된 구현이 없습니다.
Tasks
Action RecognitionResults from the Paper
| Rank | Task | Dataset | Model | Metrics |
|---|---|---|---|---|
| #4 | Action Recognition | EPIC-KITCHENS-100 | TrAction | Action@1: 54 |
| #120 | Action Recognition | Something-Something V2 | TrAction | Top-1 Accuracy: 45 |
Similar Papers 제목 키워드 기반
Coding Kendall's Shape Trajectories for 3D Action Recognition
Suitable shape representations as well as their temporal evolution, termed trajectories, often lie to non-linear manifolds. This puts an additional constraint (i.e., non-linearity) in using conventional machine learning …
3D Action RecognitionAction RecognitionDictionary LearningEvent Detection+2Hierarchical Deep Multiagent Reinforcement Learning with Temporal Abstraction
Multiagent reinforcement learning (MARL) is commonly considered to suffer from non-stationary environments and exponentially increasing policy space. It would be even more challenging when rewards are sparse and delayed …
reinforcement-learningReinforcement LearningReinforcement Learning (RL)Sparse Coding of Shape Trajectories for Facial Expression and Action Recognition
The detection and tracking of human landmarks in video streams has gained in reliability partly due to the availability of affordable RGB-D sensors. The analysis of such time-varying geometric data is playing an importan…
Action RecognitionDictionary LearningMicro Expression RecognitionMicro-Expression Recognition+2Graph Neural Network based Handwritten Trajectories Recognition
The graph neural networks has been proved to be an efficient machine learning technique in real life applications. The handwritten recognition is one of the useful area in real life use where both offline and online hand…
Graph Neural NetworkHandwriting RecognitionMotion Part Regularization: Improving Action Recognition via Trajectory Selection
Dense local motion features such as dense trajectories have been widely used in action recognition. For most actions, only a few local features (e.g., critical movements of the hand, arm, leg etc.) are responsible to the…
Action RecognitionSentenceTemporal Action Localizationtext-classification+1