paper-with-me

Papers

Technical Report: Temporal Aggregate Representations

2021-06-06 · Fadime Sener, Dibyadip Chatterjee, Angela Yao

This technical report extends our work presented in [9] with more experiments. In [9], we tackle long-term video understanding, which requires reasoning from current and past or future observations and raises several fundamental questions. How should temporal or sequential relationships be modelled? What temporal extent of information and context needs to be processed? At what temporal scale should they be derived? [9] addresses these questions with a flexible multi-granular temporal aggregation framework. In this report, we conduct further experiments with this framework on different tasks and a new dataset, EPIC-KITCHENS-100.

📄 PDF Abstract BibTeX arXiv:2106.03152

Code (1)

dibschat/tempAgg 공식 구현 pytorch

Tasks

Action AnticipationAction RecognitionVideo Understanding

Similar Papers 제목 키워드 기반

TransAction: ICL-SJTU Submission to EPIC-Kitchens Action Anticipation Challenge 2021

2021-07-28 · Xiao Gu, Jianing Qiu, Yao Guo, Benny Lo 외

In this report, the technical details of our submission to the EPIC-Kitchens Action Anticipation Challenge 2021 are given. We developed a hierarchical attention model for action anticipation, which leverages Transformer-…

Action Anticipation

CASTLE2026 Team WDL Technical Report

2026-05-30 · Zhengyang Li, Zhenglin Du, Yi Wen, Fang Liu 외 arxiv

The CASTLE Challenge @ EgoVis 2026 evaluates long-form egocentric video question answering over 600+ hours of multi-perspective recordings. Each four-choice question requires evidence from videos, transcripts, auxiliary …

Video Question AnsweringMultimodal Reasoning

Learning Effective NeRFs and SDFs Representations with 3D Generative Adversarial Networks for 3D Object Generation: Technical Report for ICCV 2023 OmniObject3D Challenge

2023-09-28 · Zheyuan Yang, Yibo Liu, Guile Wu, Tongtong Cao 외

In this technical report, we present a solution for 3D object generation of ICCV 2023 OmniObject3D Challenge. In recent years, 3D object generation has made great process and achieved promising results, but it remains a …

DecoderObject

Relation-Aware Pyramid Network (RapNet) for temporal action proposal

2019-08-09 · Jialin Gao, Zhixiang Shi, Jiani Li, Yufeng Yuan 외

In this technical report, we describe our solution to temporal action proposal (task 1) in ActivityNet Challenge 2019. First, we fine-tune a ResNet-50-C3D CNN on ActivityNet v1.3 based on Kinetics pretrained model to ext…

Relation

TGBFormer: Transformer-GraphFormer Blender Network for Video Object Detection

2025-03-18 · Qiang Qi, Xiao Wang

Video object detection has made significant progress in recent years thanks to convolutional neural networks (CNNs) and vision transformers (ViTs). Typically, CNNs excel at capturing local features but struggle to model …

GPUobject-detectionObject DetectionVideo Object Detection