Describe and Attend to Track: Learning Natural Language guided Structural Representation and Visual Attention for Object Tracking
The tracking-by-detection framework requires a set of positive and negative training samples to learn robust tracking models for precise localization of target objects. However, existing tracking models mostly treat different samples independently while ignores the relationship information among them. In this paper, we propose a novel structure-aware deep neural network to overcome such limitations. In particular, we construct a graph to represent the pairwise relationships among training samples, and additionally take the natural language as the supervised information to learn both feature representations and classifiers robustly. To refine the states of the target and re-track the target when it is back to view from heavy occlusion and out of view, we elaborately design a novel subnetwork to learn the target-driven visual attentions from the guidance of both visual and natural language cues. Extensive experiments on five tracking benchmark datasets validated the effectiveness of our proposed method.
Code (0)
등록된 구현이 없습니다.
Tasks
Object TrackingSimilar Papers 제목 키워드 기반
Eye Gaze and Self-attention: How Humans and Transformers Attend Words in Sentences
Attention describes cognitive processes that are important to many human phenomena including reading. The term is also used to describe the way in which transformer neural networks perform natural language processing. Wh…
Advances in Online Audio-Visual Meeting Transcription
This paper describes a system that generates speaker-annotated transcripts of meetings by using a microphone array and a 360-degree camera. The hallmark of the system is its ability to handle overlapped speech, which has…
Sound Source Localizationspeaker-diarizationSpeaker DiarizationSpeaker Identification+1Zero-shot Translation of Attention Patterns in VQA Models to Natural Language
Converting a model's internals to text can yield human-understandable insights about the model. Inspired by the recent success of training-free approaches for image captioning, we propose ZS-A2T, a zero-shot framework th…
Image CaptioningLanguage ModelingLanguage ModellingLarge Language Model+3ChartGaze: Enhancing Chart Understanding in LVLMs with Eye-Tracking Guided Attention Refinement
Charts are a crucial visual medium for communicating and representing information. While Large Vision-Language Models (LVLMs) have made progress on chart question answering (CQA), the task remains challenging, particular…
Chart Question AnsweringAutomated Attendee Recognition System for Large-Scale Social Events or Conference Gathering
Manual attendance tracking at large-scale events, such as marriage functions or conferences, is often inefficient and prone to human error. To address this challenge, we propose an automated, cloud-based attendance track…
Face Detection