paper-with-me

홈 › Papers

Where a Strong Backbone Meets Strong Features -- ActionFormer for Ego4D Moment Queries Challenge

2022-11-16 · Fangzhou Mu, Sicheng Mo, Gillian Wang, Yin Li

This report describes our submission to the Ego4D Moment Queries Challenge 2022. Our submission builds on ActionFormer, the state-of-the-art backbone for temporal action localization, and a trio of strong video features from SlowFast, Omnivore and EgoVLP. Our solution is ranked 2nd on the public leaderboard with 21.76% average mAP on the test set, which is nearly three times higher than the official baseline. Further, we obtain 42.54% Recall@1x at tIoU=0.5 on the test set, outperforming the top-ranked solution by a significant margin of 1.41 absolute percentage points. Our code is available at https://github.com/happyharrycn/actionformer_release.

📄 PDF Abstract BibTeX arXiv:2211.09074

Code (2)

happyharrycn/actionformer_release 공식 구현 pytorch
showlab/egovlp 공식 구현 pytorch

Tasks

Action LocalizationMoment QueriesTemporal Action Localization

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

4D Radar Meets LiDAR and Camera: Cooperative Perception under Adverse Weather

2026-05-29 · Melih Yazgan, Iramm Hamdard, Qiyuan Wu, J. Marius Zoellner arxiv

Cooperative perception is important for autonomous driving but remains fragile when cameras and LiDAR degrade in adverse weather. We address this challenge by integrating 4D imaging radar as a weather-robust modality int…

Autonomous Driving

When Shift Operation Meets Vision Transformer: An Extremely Simple Alternative to Attention Mechanism

2022-01-26 · Guangting Wang, Yucheng Zhao, Chuanxin Tang, Chong Luo 외

Attention mechanism has been widely believed as the key to success of vision transformers (ViTs), since it provides a flexible and powerful way to model spatial relationships. However, is the attention mechanism truly an…

Image ClassificationObject DetectionSemantic Segmentation

Do VLMs Need Vision Transformers? Evaluating State Space Models as Vision Encoders

2026-03-19 · Shang-Jui Ray Kuo, Paola Cascante-Bonilla arxiv

Large vision--language models (VLMs) often use a frozen vision backbone, whose image features are mapped into a large language model through a lightweight connector. While transformer-based encoders are the standard visu…

Exploring 2D backbone effects for indoor semantic occupancy prediction

2026-09-15 · Shizhang Fanga, Wanling Yea, Qi Zheng arxiv

Semantic occupancy prediction gives an embodied agent a voxel-level account of where space is free, occupied, and semantically meaningful. In RGB-D pipelines such as EmbodiedScan, the image encoder is often left as a def…

A Strong Feature Representation for Siamese Network Tracker

2019-07-18 · Zhipeng Zhou, Rui Zhang, Dong Yin

Object tracking has important application in assistive technologies for personalized monitoring. Recent trackers choosing AlexNet as their backbone to extract features have gained great success. However, AlexNet is too s…

Object Tracking