paper-with-me

Papers

Video Imprint

2021-06-07 · Zhanning Gao, Le Wang, Nebojsa Jojic, Zhenxing Niu, Nanning Zheng, Gang Hua

A new unified video analytics framework (ER3) is proposed for complex event retrieval, recognition and recounting, based on the proposed video imprint representation, which exploits temporal correlations among image features across video frames. With the video imprint representation, it is convenient to reverse map back to both temporal and spatial locations in video frames, allowing for both key frame identification and key areas localization within each frame. In the proposed framework, a dedicated feature alignment module is incorporated for redundancy removal across frames to produce the tensor representation, i.e., the video imprint. Subsequently, the video imprint is individually fed into both a reasoning network and a feature aggregation module, for event recognition/recounting and event retrieval tasks, respectively. Thanks to its attention mechanism inspired by the memory networks used in language modeling, the proposed reasoning network is capable of simultaneous event category recognition and localization of the key pieces of evidence for event recounting. In addition, the latent structure in our reasoning network highlights the areas of the video imprint, which can be directly used for event recounting. With the event retrieval task, the compact video representation aggregated from the video imprint contributes to better retrieval results than existing state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:2106.03283

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingRetrieval

Similar Papers 제목 키워드 기반

ER3: A Unified Framework for Event Retrieval, Recognition and Recounting

2017-07-01 · CVPR 2017 7 · Zhanning Gao, Gang Hua, Dong-Qing Zhang, Nebojsa Jojic 외

We develop a unified framework for complex event retrieval, recognition and recounting. The framework is based on a compact video representation that exploits the temporal correlations in image features. Our feature alig…

Language ModelingLanguage ModellingRetrieval

Revealing the Implicit Noise-based Imprint of Generative Models

2025-03-12 · Xinghan Li, Jingjing Chen, Yue Yu, Xue Song 외

With the rapid advancement of vision generation models, the potential security risks stemming from synthetic visual content have garnered increasing attention, posing significant challenges for AI-generated image detecti…

Tactile Mapping and Localization from High-Resolution Tactile Imprints

2019-04-24 · Maria Bauza, Oleguer Canal, Alberto Rodriguez

This work studies the problem of shape reconstruction and object localization using a vision-based tactile sensor, GelSlim. The main contributions are the recovery of local shapes from contact, an approach to reconstruct…

ObjectObject LocalizationVocal Bursts Intensity Prediction

Low-Shot Learning with Imprinted Weights

2017-12-19 · CVPR 2018 6 · Hang Qi, Matthew Brown, David G. Lowe

Human vision is able to immediately recognize novel visual categories after seeing just one or a few training examples. We describe how to add a similar capability to ConvNet classifiers by directly setting the final lay…

From Pixels to Trajectory: Universal Adversarial Example Detection via Temporal Imprints

2025-03-06 · Yansong Gao, Huaibing Peng, Hua Ma, Zhiyang Dai 외

For the first time, we unveil discernible temporal (or historical) trajectory imprints resulting from adversarial example (AE) attacks. Standing in contrast to existing studies all focusing on spatial (or static) imprint…

One-Class Classification