paper-with-me

Papers

BUS:Efficient and Effective Vision-language Pre-training with Bottom-Up Patch Summarization

2023-07-17 · Chaoya Jiang, Haiyang Xu, Wei Ye, Qinghao Ye, Chenliang Li, Ming Yan, Bin Bi, Shikun Zhang, Fei Huang, Songfang Huang

Vision Transformer (ViT) based Vision-Language Pre-training (VLP) models have demonstrated impressive performance in various tasks. However, the lengthy visual token sequences fed into ViT can lead to training inefficiency and ineffectiveness. Existing efforts address the challenge by either bottom-level patch extraction in the ViT backbone or top-level patch abstraction outside, not balancing training efficiency and effectiveness well. Inspired by text summarization in natural language processing, we propose a Bottom-Up Patch Summarization approach named BUS, coordinating bottom-level extraction and top-level abstraction to learn a concise summary of lengthy visual token sequences efficiently. Specifically, We incorporate a Text-Semantics-Aware Patch Selector (TSPS) into the ViT backbone to perform a coarse-grained visual token extraction and then attach a flexible Transformer-based Patch Abstraction Decoder (PAD) upon the backbone for top-level visual abstraction. This bottom-up collaboration enables our BUS to yield high training efficiency while maintaining or even improving effectiveness. We evaluate our approach on various visual-language understanding and generation tasks and show competitive downstream task performance while boosting the training efficiency by 50\%. Additionally, our model achieves state-of-the-art performance on many downstream tasks by increasing input image resolution without increasing computational costs over baselines.

📄 PDF Abstract BibTeX arXiv:2307.08504

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderText Summarization

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Position-Wise Feed-Forward Layer 설명 없음
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Residual Connection 설명 없음
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…

Similar Papers 제목 키워드 기반

BUS: Efficient and Effective Vision-Language Pre-Training with Bottom-Up Patch Summarization.

2023-01-01 · ICCV 2023 1 · Chaoya Jiang, Haiyang Xu, Wei Ye, Qinghao Ye 외

Vision Transformer (ViT) based Vision-Language Pretraining (VLP) models recently demonstrated impressive performance in various tasks. However, the lengthy visual token sequences used in these models can lead to inef…

Abstractive Text SummarizationDecoderDocument Summarization

Vision-LSTM: xLSTM as Generic Vision Backbone

2024-06-06 · Benedikt Alkin, Maximilian Beck, Korbinian Pöppel, Sepp Hochreiter 외

Transformers are widely used as generic backbones in computer vision, despite initially introduced for natural language processing. Recently, the Long Short-Term Memory (LSTM) has been extended to a scalable and performa…

ScanFormer: Referring Expression Comprehension by Iteratively Scanning

2024-06-26 · CVPR 2024 1 · Wei Su, Peihan Miao, Huanzhang Dou, Xi Li

Referring Expression Comprehension (REC) aims to localize the target objects specified by free-form natural language descriptions in images. While state-of-the-art methods achieve impressive performance, they perform a d…

InformativenessReferring ExpressionReferring Expression Comprehension

HEED: Density-Weighted Residual Alignment for Hybrid Vision-Language Model Distillation

2026-05-16 · Yihao Liang, Niraj K. Jha arxiv

Distilling vision-language models into faster hybrid architectures, such as 3:1 Mamba-2/attention mixes, is now standard practice for making inference efficient. Aggregate benchmarks suggest that this works but they hide…

Visual Reasoning

Distilling semantically aware orders for autoregressive image generation

2025-04-23 · Rishav Pramanik, Antoine Poupon, Juan A. Rodriguez, Masih Aminbeidokhti 외

Autoregressive patch-based image generation has recently shown competitive results in terms of image quality and scalability. It can also be easily integrated and scaled within Vision-Language models. Nevertheless, autor…

Image GenerationText GenerationText-to-Image Generation