paper-with-me

홈 › Papers

Ladder Loss for Coherent Visual-Semantic Embedding

2019-11-18 · Mo Zhou, Zhenxing Niu, Le Wang, Zhanning Gao, Qilin Zhang, Gang Hua

For visual-semantic embedding, the existing methods normally treat the relevance between queries and candidates in a bipolar way -- relevant or irrelevant, and all "irrelevant" candidates are uniformly pushed away from the query by an equal margin in the embedding space, regardless of their various proximity to the query. This practice disregards relatively discriminative information and could lead to suboptimal ranking in the retrieval results and poorer user experience, especially in the long-tail query scenario where a matching candidate may not necessarily exist. In this paper, we introduce a continuous variable to model the relevance degree between queries and multiple candidates, and propose to learn a coherent embedding space, where candidates with higher relevance degrees are mapped closer to the query than those with lower relevance degrees. In particular, the new ladder loss is proposed by extending the triplet loss inequality to a more general inequality chain, which implements variable push-away margins according to respective relevance degrees. In addition, a proper Coherent Score metric is proposed to better measure the ranking results including those "irrelevant" candidates. Extensive experiments on multiple datasets validate the efficacy of our proposed method, which achieves significant improvement over existing state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:1911.07528

Code (2)

cdluminate/ladderloss 공식 구현 pytorch
cdluminate/cdluminate

Tasks

RetrievalTriplet

Methods 이 논문이 사용한 방법론

Triplet Loss The goal of Triplet loss, in the context of Siamese Networks, is to maximize the joint probability among all score-pairs i.e. the product of all probabilities. By using its…

Similar Papers 제목 키워드 기반

Mathematical Justification of Hard Negative Mining via Isometric Approximation Theorem

2022-10-20 · Albert Xu, Jhih-Yi Hsieh, Bhaskar Vundurthy, Eliana Cohen 외

In deep metric learning, the Triplet Loss has emerged as a popular method to learn many computer vision and natural language processing tasks such as facial recognition, object detection, and visual-semantic embeddings. …

Contrastive LearningMetric Learningobject-detectionObject Detection+1

A Quadruplet Loss for Enforcing Semantically Coherent Embeddings in Multi-output Classification Problems

2020-02-26 · Hugo Proença, Ehsan Yaghoubi, Pendar Alirezazadeh

This paper describes one objective function for learning semantically coherent feature embeddings in multi-output classification problems, i.e., when the response variables have dimension higher than one. In particular, …

General ClassificationRetrievalSemantic SimilaritySemantic Textual Similarity+1

VLTinT: Visual-Linguistic Transformer-in-Transformer for Coherent Video Paragraph Captioning

2022-11-28 · Kashu Yamazaki, Khoa Vo, Sang Truong, Bhiksha Raj 외

Video paragraph captioning aims to generate a multi-sentence description of an untrimmed video with several temporal event locations in coherent storytelling. Following the human perception process, where the scene is ef…

DiversitySentenceVideo Captioning

The Semantic Ladder: A Framework for Progressive Formalization of Natural Language Content for Knowledge Graphs and AI Systems

2026-03-23 · Lars Vogt arxiv

Semantic data and knowledge infrastructures must reconcile two fundamentally different forms of representation: natural language, in which most knowledge is created and communicated, and formal semantic models, which ena…

Knowledge GraphsSemantic Parsing

Continual Learning for Image Captioning through Improved Image-Text Alignment

2025-10-07 · Bertram Taetz, Gal Bordelius arxiv

Generating accurate and coherent image captions in a continual learning setting remains a major challenge due to catastrophic forgetting and the difficulty of aligning evolving visual concepts with language over time. In…

Continual LearningImage Captioning