paper-with-me

홈 › Papers

GLAD: Generative Language-Assisted Visual Tracking for Low-Semantic Templates

2026-01-31 · Xingyu Luo, Yidong Cai, Jie Liu, Jie Tang, Gangshan Wu, Limin Wang arxiv

Vision-language tracking has gained increasing attention in many scenarios. This task simultaneously deals with visual and linguistic information to localize objects in videos. Despite its growing utility, the development of vision-language tracking methods remains in its early stage. Current vision-language trackers usually employ Transformer architectures for interactive integration of template, search, and text features. However, persistent challenges about low-semantic images including prevalent image blurriness, low resolution and so on, may compromise model performance through degraded cross-modal understanding. To solve this problem, language assistance is usually used to deal with the obstacles posed by low-semantic images. However, due to the existing gap between current textual and visual features, direct concatenation and fusion of these features may have limited effectiveness. To address these challenges, we introduce a pioneering Generative Language-AssisteD tracking model, GLAD, which utilizes diffusion models for the generative multi-modal fusion of text description and template image to bolster compatibility between language and image and enhance template image semantic information. Our approach demonstrates notable improvements over the existing fusion paradigms. Blurry and semantically ambiguous template images can be restored to improve multi-modal features in the generative fusion paradigm. Experiments show that our method establishes a new state-of-the-art on multiple benchmarks and achieves an impressive inference speed. The code and models will be released at: https://github.com/Confetti-lxy/GLAD

📄 PDF Abstract BibTeX arXiv:2602.00570

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Tracking

Similar Papers 제목 키워드 기반

Oitijjo-3D: Generative AI Framework for Rapid 3D Heritage Reconstruction from Street View Imagery

2025-11-01 · Momen Khandoker Ope, Akif Islam, Mohd Ruhul Ameen, Abu Saleh Musa Miah 외 arxiv

Cultural heritage restoration in Bangladesh faces a dual challenge of limited resources and scarce technical expertise. Traditional 3D digitization methods, such as photogrammetry or LiDAR scanning, require expensive har…

Visual Reasoning3D Generation

The differentials and determinants of perinatal mortality in rural Bangladesh

2003-01-01 · journal 2003 1 · W. Bari, R. I. Chowdhury, M. A. Islam, N. Chakraborty 외

Objective In Bangladesh, the perinatal mortality is very high. This study examined the differentials and determinants of perinatal mortality in rural Bangladesh. Methods The study was based on the prospective data on m…

Toward Scalable Neural Dialogue State Tracking Model

2018-12-03 · Elnaz Nouri, Ehsan Hosseini-Asl

The latency in the current neural based dialogue state tracking models prohibits them from being used efficiently for deployment in production systems, albeit their highly accurate performance. This paper proposes a new …

Dialogue State TrackingmodelMulti-domain Dialogue State Tracking

Mind the Gap: Geometrically Accurate Generative Reconstruction from Disjoint Views

2026-05-08 · Grzegorz Wilczynski, Mikołaj Zielinski, Bartosz Świrta, Dominik Belter 외 arxiv

3D vision systems are fundamentally constrained by their reliance on visual overlap: reconstruction methods require it for geometric alignment, while generative models use it to enforce multi-view consistency. This limit…

3D Reconstruction

Global-Locally Self-Attentive Encoder for Dialogue State Tracking

2018-07-01 · ACL 2018 7 · Victor Zhong, Caiming Xiong, Richard Socher

Dialogue state tracking, which estimates user goals and requests given the dialogue context, is an essential part of task-oriented dialogue systems. In this paper, we propose the Global-Locally Self-Attentive Dialogue St…

Automatic Speech Recognition (ASR)Dialogue State TrackingRepresentation LearningSpeech Recognition+2