paper-with-me

홈 › Papers

Visual Transformer for Task-aware Active Learning

2021-06-07 · Razvan Caramalau, Binod Bhattarai, Tae-Kyun Kim

Pool-based sampling in active learning (AL) represents a key framework for an-notating informative data when dealing with deep learning models. In this paper, we present a novel pipeline for pool-based Active Learning. Unlike most previous works, our method exploits accessible unlabelled examples during training to estimate their co-relation with the labelled examples. Another contribution of this paper is to adapt Visual Transformer as a sampler in the AL pipeline. Visual Transformer models non-local visual concept dependency between labelled and unlabelled examples, which is crucial to identifying the influencing unlabelled examples. Also, compared to existing methods where the learner and the sampler are trained in a multi-stage manner, we propose to train them in a task-aware jointly manner which enables transforming the latent space into two separate tasks: one that classifies the labelled examples; the other that distinguishes the labelling direction. We evaluated our work on four different challenging benchmarks of classification and detection tasks viz. CIFAR10, CIFAR100,FashionMNIST, RaFD, and Pascal VOC 2007. Our extensive empirical and qualitative evaluations demonstrate the superiority of our method compared to the existing methods. Code available: https://github.com/razvancaramalau/Visual-Transformer-for-Task-aware-Active-Learning

📄 PDF Abstract BibTeX arXiv:2106.03801

Code (1)

razvancaramalau/Visual-Transformer-for-Task-aware-Active-Learning 공식 구현 pytorch

Tasks

Active Learning

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Position-Wise Feed-Forward Layer 설명 없음
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Adam 설명 없음
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…

Similar Papers 제목 키워드 기반

TransRefer3D: Entity-and-Relation Aware Transformer for Fine-Grained 3D Visual Grounding

2021-08-05 · Dailan He, Yusheng Zhao, Junyu Luo, Tianrui Hui 외

Recently proposed fine-grained 3D visual grounding is an essential and challenging task, whose goal is to identify the 3D object referred by a natural language sentence from other distractive objects of the same category…

3D visual groundingRelationSentenceVisual Grounding

Stepwise Extractive Summarization and Planning with Structured Transformers

2020-10-06 · EMNLP 2020 11 · Shashi Narayan, Joshua Maynez, Jakub Adamek, Daniele Pighin 외

We propose encoder-centric stepwise models for extractive summarization using structured transformers -- HiBERT and Extended Transformers. We enable stepwise summarization by injecting the previously generated summary in…

Extractive SummarizationSentenceTable-to-Text GenerationText Generation

Set2Seq Transformer: Learning Permutation Aware Set Representations of Artistic Sequences

2024-08-06 · Athanasios Efthymiou, Stevan Rudinac, Monika Kackovic, Nachoem Wijnberg 외

We propose Set2Seq Transformer, a novel sequential multiple instance architecture, that learns to rank permutation aware set representations of sequences. First, we illustrate that learning temporal position-aware repres…

Art AnalysisLearning-To-RankMultiple Instance LearningPosition

HSVLT: Hierarchical Scale-Aware Vision-Language Transformer for Multi-Label Image Classification

2024-07-23 · Shuyi Ouyang, Hongyi Wang, Ziwei Niu, Zhenjia Bai 외

The task of multi-label image classification involves recognizing multiple objects within a single image. Considering both valuable semantic information contained in the labels and essential visual features presented in …

image-classificationImage ClassificationMulti-Label Image Classification

Vision And Text Transformer For Predicting Answerability On Visual Question Answering

2026-09-15 · Tung Le, Huy Tien Nguyen, Le Minh Nguyen arxiv

Answerability on Visual Question Answering is a novel and attractive task to predict answerable scores between images and questions in multi-modal data. Existing works often utilize a binary mapping from visual question …

Visual Question Answering