paper-with-me

홈 › Papers

Describe and Attend to Track: Learning Natural Language guided Structural Representation and Visual Attention for Object Tracking

2018-11-25 · Xiao Wang, Chenglong Li, Rui Yang, Tianzhu Zhang, Jin Tang, Bin Luo

The tracking-by-detection framework requires a set of positive and negative training samples to learn robust tracking models for precise localization of target objects. However, existing tracking models mostly treat different samples independently while ignores the relationship information among them. In this paper, we propose a novel structure-aware deep neural network to overcome such limitations. In particular, we construct a graph to represent the pairwise relationships among training samples, and additionally take the natural language as the supervised information to learn both feature representations and classifiers robustly. To refine the states of the target and re-track the target when it is back to view from heavy occlusion and out of view, we elaborately design a novel subnetwork to learn the target-driven visual attentions from the guidance of both visual and natural language cues. Extensive experiments on five tracking benchmark datasets validated the effectiveness of our proposed method.

📄 PDF Abstract BibTeX arXiv:1811.10014

Code (0)

등록된 구현이 없습니다.

Tasks

Object Tracking

Similar Papers 제목 키워드 기반

Eye Gaze and Self-attention: How Humans and Transformers Attend Words in Sentences

2022-05-01 · CMCL (ACL) 2022 5 · Joshua Bensemann, Alex Peng, Diana Prado, Yang Chen 외

Attention describes cognitive processes that are important to many human phenomena including reading. The term is also used to describe the way in which transformer neural networks perform natural language processing. Wh…

Advances in Online Audio-Visual Meeting Transcription

2019-12-10 · Takuya Yoshioka, Igor Abramovski, Cem Aksoylar, Zhuo Chen 외

This paper describes a system that generates speaker-annotated transcripts of meetings by using a microphone array and a 360-degree camera. The hallmark of the system is its ability to handle overlapped speech, which has…

Sound Source Localizationspeaker-diarizationSpeaker DiarizationSpeaker Identification+1

Zero-shot Translation of Attention Patterns in VQA Models to Natural Language

2023-11-08 · Leonard Salewski, A. Sophia Koepke, Hendrik P. A. Lensch, Zeynep Akata

Converting a model's internals to text can yield human-understandable insights about the model. Inspired by the recent success of training-free approaches for image captioning, we propose ZS-A2T, a zero-shot framework th…

Image CaptioningLanguage ModelingLanguage ModellingLarge Language Model+3

ChartGaze: Enhancing Chart Understanding in LVLMs with Eye-Tracking Guided Attention Refinement

2025-09-16 · Ali Salamatian, Amirhossein Abaskohi, Wan-Cyuan Fan, Mir Rayat Imtiaz Hossain 외 arxiv

Charts are a crucial visual medium for communicating and representing information. While Large Vision-Language Models (LVLMs) have made progress on chart question answering (CQA), the task remains challenging, particular…

Chart Question Answering

Automated Attendee Recognition System for Large-Scale Social Events or Conference Gathering

2025-03-05 · Dhruv Motwani, Ankush Tyagi, Vipul Dabhi, Harshadkumar Prajapati

Manual attendance tracking at large-scale events, such as marriage functions or conferences, is often inefficient and prone to human error. To address this challenge, we propose an automated, cloud-based attendance track…

Face Detection