paper-with-me

홈 › Papers

SiRi: A Simple Selective Retraining Mechanism for Transformer-based Visual Grounding

2022-07-27 · Mengxue Qu, Yu Wu, Wu Liu, Qiqi Gong, Xiaodan Liang, Olga Russakovsky, Yao Zhao, Yunchao Wei

In this paper, we investigate how to achieve better visual grounding with modern vision-language transformers, and propose a simple yet powerful Selective Retraining (SiRi) mechanism for this challenging task. Particularly, SiRi conveys a significant principle to the research of visual grounding, i.e., a better initialized vision-language encoder would help the model converge to a better local minimum, advancing the performance accordingly. In specific, we continually update the parameters of the encoder as the training goes on, while periodically re-initialize rest of the parameters to compel the model to be better optimized based on an enhanced encoder. SiRi can significantly outperform previous approaches on three popular benchmarks. Specifically, our method achieves 83.04% Top1 accuracy on RefCOCO+ testA, outperforming the state-of-the-art approaches (training from scratch) by more than 10.21%. Additionally, we reveal that SiRi performs surprisingly superior even with limited training data. We also extend it to transformer-based visual grounding models and other vision-language tasks to verify the validity.

📄 PDF Abstract BibTeX arXiv:2207.13325

Code (1)

qumengxue/siri-vg 공식 구현 pytorch

Tasks

Visual Grounding

Similar Papers 제목 키워드 기반

Selective Attention Improves Transformer

2024-10-03 · Yaniv Leviathan, Matan Kalman, Yossi Matias

Unneeded elements in the attention's context degrade performance. We introduce Selective Attention, a simple parameter-free change to the standard attention mechanism which reduces attention to unneeded elements. Selecti…

Language ModelingLanguage Modelling

Symmetry Breaking in Transformers for Efficient and Interpretable Training

2026-01-29 · Eva Silverstein, Daniel Kunin, Vasudev Shyam arxiv

The attention mechanism in its standard implementation contains extraneous rotational degrees of freedom that are carried through computation but do not affect model activations or outputs. We introduce a simple symmetry…

Logical Reasoning

Masked Retraining Teacher-Student Framework for Domain Adaptive Object Detection

2023-01-01 · ICCV 2023 1 · Zijing Zhao, Sitong Wei, Qingchao Chen, Dehui Li 외

Domain adaptive Object Detection (DAOD) leverages a labeled domain (source) to learn an object detector generalizing to a novel domain without annotation (target). Recent advances use a teacher-student framework, i.e…

Decoderobject-detectionObject DetectionUnsupervised Domain Adaptation

Sirius: Contextual Sparsity with Correction for Efficient LLMs

2024-09-05 · Yang Zhou, Zhuoming Chen, Zhaozhuo Xu, Victoria Lin 외

With the blossom of large language models (LLMs), inference efficiency becomes increasingly important. Various approximation methods are proposed to reduce the cost at inference time. Contextual Sparsity (CS) is appealin…

Math

Fusion: A Framework for Unified Sequential Token AdaptatIon in VisiOn TraNsformers

2026-07-01 · Aravind Pradeep, Samira Nazari, Mahdi Taheri, Christian Herglotz arxiv

Vision Transformers achieve strong image classification accuracy but process all image regions with nearly the same computation, even when many regions are redundant or uninformative. Recent adaptive inference methods re…

Image Classification