paper-with-me

홈 › Papers

Adaptive Offline Quintuplet Loss for Image-Text Matching

2020-03-07 · ECCV 2020 8 · Tianlang Chen, Jiajun Deng, Jiebo Luo

Existing image-text matching approaches typically leverage triplet loss with online hard negatives to train the model. For each image or text anchor in a training mini-batch, the model is trained to distinguish between a positive and the most confusing negative of the anchor mined from the mini-batch (i.e. online hard negative). This strategy improves the model's capacity to discover fine-grained correspondences and non-correspondences between image and text inputs. However, the above approach has the following drawbacks: (1) the negative selection strategy still provides limited chances for the model to learn from very hard-to-distinguish cases. (2) The trained model has weak generalization capability from the training set to the testing set. (3) The penalty lacks hierarchy and adaptiveness for hard negatives with different "hardness" degrees. In this paper, we propose solutions by sampling negatives offline from the whole training set. It provides "harder" offline negatives than online hard negatives for the model to distinguish. Based on the offline hard negatives, a quintuplet loss is proposed to improve the model's generalization capability to distinguish positives and negatives. In addition, a novel loss function that combines the knowledge of positives, offline hard negatives and online hard negatives is created. It leverages offline hard negatives as the intermediary to adaptively penalize them based on their distance relations to the anchor. We evaluate the proposed training approach on three state-of-the-art image-text models on the MS-COCO and Flickr30K datasets. Significant performance improvements are observed for all the models, proving the effectiveness and generality of our approach. Code is available at https://github.com/sunnychencool/AOQ

📄 PDF Abstract BibTeX arXiv:2003.03669

Code (1)

sunnychencool/AOQ 공식 구현 pytorch

Tasks

Image-text matchingText MatchingTriplet

Methods 이 논문이 사용한 방법론

Triplet Loss The goal of Triplet loss, in the context of Siamese Networks, is to maximize the joint probability among all score-pairs i.e. the product of all probabilities. By using its…

Similar Papers 제목 키워드 기반

Learning Quintuplet Loss for Large-scale Visual Geo-Localization

2019-07-26 · Qiang Zhai

With the maturity of Artificial Intelligence (AI) technology, Large Scale Visual Geo-Localization (LSVGL) is increasingly important in urban computing, where the task is to accurately and efficiently recognize the geo-lo…

geo-localizationMetric LearningTripletVisual Place Recognition

Learning Joint Gait Representation via Quintuplet Loss Minimization

2019-06-01 · CVPR 2019 6 · Kaihao Zhang, Wenhan Luo, Lin Ma, Wei Liu 외

Gait recognition is an important biometric method popularly used in video surveillance, where the task is to identify people at a distance by their walking patterns from video sequences. Most of the current successful ap…

Gait Recognition

TRACE: A Differentiable Approach to Line-level Stroke Recovery for Offline Handwritten Text

2021-05-24 · Taylor Archibald, Mason Poggemann, Aaron Chan, Tony Martinez

Stroke order and velocity are helpful features in the fields of signature verification, handwriting recognition, and handwriting synthesis. Recovering these features from offline handwritten text is a challenging and wel…

Dynamic Time WarpingHandwriting RecognitionTrajectory Recovery

Adaptive Behavior Cloning Regularization for Stable Offline-to-Online Reinforcement Learning

2022-10-25 · Yi Zhao, Rinu Boney, Alexander Ilin, Juho Kannala 외

Offline reinforcement learning, by learning from a fixed dataset, makes it possible to learn agent behaviors without interacting with the environment. However, depending on the quality of the offline dataset, such pre-tr…

D4RLOffline RLreinforcement-learningReinforcement Learning+1

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models

2025-09-04 · Hongyin Zhang, Shiyuan Zhang, Junxi Jin, Qixin Zeng 외 arxiv

Vision-Language-Action (VLA) models based on flow matching have shown excellent performance in general-purpose robotic manipulation tasks. However, the action accuracy of these models on complex downstream tasks is unsat…

Reinforcement LearningFew-Shot LearningOffline RL