paper-with-me

Papers

Beyond Visual Cues: Synchronously Exploring Target-Centric Semantics for Vision-Language Tracking

2023-11-28 · Jiawei Ge, Xiangmei Chen, Jiuxin Cao, Xuelin Zhu, Bo Liu

Single object tracking aims to locate one specific target in video sequences, given its initial state. Classical trackers rely solely on visual cues, restricting their ability to handle challenges such as appearance variations, ambiguity, and distractions. Hence, Vision-Language (VL) tracking has emerged as a promising approach, incorporating language descriptions to directly provide high-level semantics and enhance tracking performance. However, current VL trackers have not fully exploited the power of VL learning, as they suffer from limitations such as heavily relying on off-the-shelf backbones for feature extraction, ineffective VL fusion designs, and the absence of VL-related loss functions. Consequently, we present a novel tracker that progressively explores target-centric semantics for VL tracking. Specifically, we propose the first Synchronous Learning Backbone (SLB) for VL tracking, which consists of two novel modules: the Target Enhance Module (TEM) and the Semantic Aware Module (SAM). These modules enable the tracker to perceive target-related semantics and comprehend the context of both visual and textual modalities at the same pace, facilitating VL feature extraction and fusion at different semantic levels. Moreover, we devise the dense matching loss to further strengthen multi-modal representation learning. Extensive experiments on VL tracking datasets demonstrate the superiority and effectiveness of our methods.

📄 PDF Abstract BibTeX arXiv:2311.17085

Code (0)

등록된 구현이 없습니다.

Tasks

Object TrackingRepresentation Learning

Methods 이 논문이 사용한 방법론

AWARE We propose to theoretically and empirically examine the effect of incorporating weighting schemes into walk-aggregating GNNs. To this end, we propose a simple, interpretable, and…

Similar Papers 제목 키워드 기반

SmartPhone: Exploring Keyword Mnemonic with Auto-generated Verbal and Visual Cues

2023-05-11 · Jaewook Lee, Andrew Lan

In second language vocabulary learning, existing works have primarily focused on either the learning interface or scheduling personalized retrieval practices to maximize memory retention. However, the learning content, i…

RetrievalScheduling

Ref-Adv: Exploring MLLM Visual Reasoning in Referring Expression Tasks

2026-02-27 · Qihua Dong, Kuo Yang, Lin Ju, Handong Zhao 외 arxiv

Referring Expression Comprehension (REC) links language to region level visual perception. Standard benchmarks (RefCOCO, RefCOCO+, RefCOCOg) have progressed rapidly with multimodal LLMs but remain weak tests of visual re…

Referring ExpressionVisual Reasoning

Beyond Visual Semantics: Exploring the Role of Scene Text in Image Understanding

2019-05-25 · Arka Ujjal Dey, Suman Kumar Ghosh, Ernest Valveny, Gaurav Harit

Images with visual and scene text content are ubiquitous in everyday life. However, current image interpretation systems are mostly limited to using only the visual features, neglecting to leverage the scene text content…

Retrieval

Exploring Image Representation with Decoupled Classical Visual Descriptors

2025-10-16 · Chenyuan Qu, Hao Chen, Jianbo Jiao arxiv

Exploring and understanding efficient image representations is a long-standing challenge in computer vision. While deep learning has achieved remarkable progress across image understanding tasks, its internal representat…

Image Generation

Beyond the Mouth: Upper-Face Affective Cues in Audiovisual Sentence Recognition under Acoustic Uncertainty

2026-05-30 · Zhou Yang, Yueyi Yang arxiv

Face-to-face speech comprehension is inherently multimodal, integrating acoustic signals with visible articulation, facial expression, head motion, and other socially relevant cues. While audiovisual speech systems typic…