paper-with-me

홈 › Papers

Identity-aware Graph Memory Network for Action Detection

2021-08-26 · Jingcheng Ni, Jie Qin, Di Huang

Action detection plays an important role in high-level video understanding and media interpretation. Many existing studies fulfill this spatio-temporal localization by modeling the context, capturing the relationship of actors, objects, and scenes conveyed in the video. However, they often universally treat all the actors without considering the consistency and distinctness between individuals, leaving much room for improvement. In this paper, we explicitly highlight the identity information of the actors in terms of both long-term and short-term context through a graph memory network, namely identity-aware graph memory network (IGMN). Specifically, we propose the hierarchical graph neural network (HGNN) to comprehensively conduct long-term relation modeling within the same identity as well as between different ones. Regarding short-term context, we develop a dual attention module (DAM) to generate identity-aware constraint to reduce the influence of interference by the actors of different identities. Extensive experiments on the challenging AVA dataset demonstrate the effectiveness of our method, which achieves state-of-the-art results on AVA v2.1 and v2.2.

📄 PDF Abstract BibTeX arXiv:2108.11559

Code (0)

등록된 구현이 없습니다.

Tasks

Action DetectionGraph Neural NetworkTemporal LocalizationVideo Understanding

Methods 이 논문이 사용한 방법론

Graph Neural Network 설명 없음
Memory Network 설명 없음

Similar Papers 제목 키워드 기반

Significant Other AI: Identity, Memory, and Emotional Regulation as Long-Term Relational Intelligence

2025-11-29 · Sung Park arxiv

Significant Others (SOs) stabilize identity, regulate emotion, and support narrative meaning-making, yet many people today lack access to such relational anchors. Recent advances in large language models and memory-augme…

MEMO: Memory-Guided Diffusion for Expressive Talking Video Generation

2024-12-05 · Longtao Zheng, Yifan Zhang, Hanzhong Guo, Jiachun Pan 외

Recent advances in video diffusion models have unlocked new potential for realistic audio-driven talking video generation. However, achieving seamless audio-lip synchronization, maintaining long-term identity consistency…

Portrait AnimationVideo Generation

UTM: A Unified Multiple Object Tracking Model With Identity-Aware Feature Enhancement

2023-01-01 · CVPR 2023 1 · Sisi You, Hantao Yao, Bing-Kun Bao, Changsheng Xu

Recently, Multiple Object Tracking has achieved great success, which consists of object detection, feature embedding, and identity association. Existing methods apply the three-step or two-step paradigm to generate r…

Multiple Object Trackingobject-detectionObject DetectionObject Tracking

ID-RAG: Identity Retrieval-Augmented Generation for Long-Horizon Persona Coherence in Generative Agents

2025-09-29 · Daniel Platnick, Mohamed E. Bengueddache, Marjan Alirezaie, Dava J. Newman 외 arxiv

Generative agents powered by language models are increasingly deployed for long-horizon tasks. However, as long-term memory context grows over time, they struggle to maintain coherence. This deficiency leads to critical …

LT-Mem: Volatility-Aware Spatio-Temporal Memory for Lifelong Scene Understanding

2026-08-19 · Yumin Lee, Hyoseok Ju, Giseop Kim arxiv

Long-term robot operation in evolving environments requires object-level understanding that persists across repeated revisits. Existing systems either overwrite history to maintain an up-to-date map or store semantic sna…

Scene Understanding