paper-with-me

홈 › Papers

Augmented 2D-TAN: A Two-stage Approach for Human-centric Spatio-Temporal Video Grounding

2021-06-20 · Chaolei Tan, Zihang Lin, Jian-Fang Hu, Xiang Li, Wei-Shi Zheng

We propose an effective two-stage approach to tackle the problem of language-based Human-centric Spatio-Temporal Video Grounding (HC-STVG) task. In the first stage, we propose an Augmented 2D Temporal Adjacent Network (Augmented 2D-TAN) to temporally ground the target moment corresponding to the given description. Primarily, we improve the original 2D-TAN from two aspects: First, a temporal context-aware Bi-LSTM Aggregation Module is developed to aggregate clip-level representations, replacing the original max-pooling. Second, we propose to employ Random Concatenation Augmentation (RCA) mechanism during the training phase. In the second stage, we use pretrained MDETR model to generate per-frame bounding boxes via language query, and design a set of hand-crafted rules to select the best matching bounding box outputted by MDETR for each frame within the grounded moment.

📄 PDF Abstract BibTeX arXiv:2106.10634

Code (0)

등록된 구현이 없습니다.

Tasks

Spatio-Temporal Video GroundingVideo Grounding

Methods 이 논문이 사용한 방법론

MDETR MDETR is an end-to-end modulated detector that detects objects in an image conditioned on a raw text query, like a caption or a question. It utilizes a…

Similar Papers 제목 키워드 기반

Spatio-temporal dual-stage hypergraph MARL for human-centric multimodal corridor traffic signal control

2026-02-19 · Xiaocai Zhang, Neema Nassir, Milad Haghani arxiv

Human-centric traffic signal control in corridor networks must increasingly account for multimodal travelers, particularly high-occupancy public transportation, rather than focusing solely on vehicle-centric performance.…

Multi-agent Reinforcement Learning

DETACH : Decomposed Spatio-Temporal Alignment for Exocentric Video and Ambient Sensors with Staged Learning

2025-12-23 · Junho Yoon, Jaemo Jung, Hyunju Kim, Dongman Lee arxiv

Aligning egocentric video with wearable sensors have shown promise for human action recognition, but face practical limitations in user discomfort, privacy concerns, and scalability. We explore exocentric video with ambi…

Action RecognitionOnline Clustering

Fine-grained Spatiotemporal Grounding on Egocentric Videos

2025-08-01 · Shuo Liang, Yiwu Zhong, Zi-Yuan Hu, Yeyao Tao 외 arxiv

Spatiotemporal video grounding aims to localize target entities in videos based on textual queries. While existing research has made significant progress in exocentric videos, the egocentric setting remains relatively un…

Video Grounding

EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT

2025-10-27 · Baoqi Pei, Yifei Huang, Jilan Xu, Yuping He 외 arxiv

Egocentric video reasoning centers on an unobservable agent behind the camera who dynamically shapes the environment, requiring inference of hidden intentions and recognition of fine-grained interactions. This core chall…

Episodic Memory Question Answering

2022-05-03 · CVPR 2022 1 · Samyak Datta, Sameer Dharur, Vincent Cartillier, Ruta Desai 외

Egocentric augmented reality devices such as wearable glasses passively capture visual data as a human wearer tours a home environment. We envision a scenario wherein the human communicates with an AI agent powering such…

AI AgentQuestion Answering