paper-with-me

Papers

Learnable PINs: Cross-Modal Embeddings for Person Identity

2018-05-02 · ECCV 2018 9 · Arsha Nagrani, Samuel Albanie, Andrew Zisserman

We propose and investigate an identity sensitive joint embedding of face and voice. Such an embedding enables cross-modal retrieval from voice to face and from face to voice. We make the following four contributions: first, we show that the embedding can be learnt from videos of talking faces, without requiring any identity labels, using a form of cross-modal self-supervision; second, we develop a curriculum learning schedule for hard negative mining targeted to this task, that is essential for learning to proceed successfully; third, we demonstrate and evaluate cross-modal retrieval for identities unseen and unheard during training over a number of scenarios and establish a benchmark for this novel task; finally, we show an application of using the joint embedding for automatically retrieving and labelling characters in TV dramas.

📄 PDF Abstract BibTeX arXiv:1805.00833

Code (1)

my-yy/learnable_pins pytorch

Tasks

Cross-Modal RetrievalRetrieval

Similar Papers 제목 키워드 기반

UEmbed: Unified Sparse and Dense Multimodal Embeddings

2026-08-03 · Tingyu Song, Mingxin Li, Yanzhao Zhang, Dingkun Long 외 hf

Sparse retrieval underpins modern search systems, from web search to retrieval-augmented generation. Existing work has introduced Learned Sparse Retrieval (LSR) to push beyond exact lexical matching toward richer semanti…

Decoupled Entity Representation Learning for Pinterest Ads Ranking

2025-09-04 · Jie Liu, Yinrui Li, Jiankai Sun, Kungang Li 외 arxiv

In this paper, we introduce a novel framework following an upstream-downstream paradigm to construct user and item (Pin) embeddings from diverse data sources, which are essential for Pinterest to deliver personalized Pin…

Representation Learning

ToPT: Task-Oriented Prompt Tuning for Urban Region Representation Learning

2026-02-02 · Zitao Guo, Changyang Jiang, Tianhong Zhao, Jinzhou Cao 외 arxiv

Learning effective region embeddings from heterogeneous urban data underpins key urban computing tasks (e.g., crime prediction, resource allocation). However, prevailing two-stage methods yield task-agnostic representati…

Representation Learning

PinnerSage: Multi-Modal User Embedding Framework for Recommendations at Pinterest

2020-07-07 · Aditya Pal, Chantat Eksombatchai, Yitong Zhou, Bo Zhao 외

Latent user representations are widely adopted in the tech industry for powering personalized recommender systems. Most prior work infers a single high dimensional embedding to represent a user, which is a good starting …

ClusteringRecommendation Systems

X-MIC: Cross-Modal Instance Conditioning for Egocentric Action Generalization

2024-03-28 · CVPR 2024 1 · Anna Kukleva, Fadime Sener, Edoardo Remelli, Bugra Tekin 외

Lately, there has been growing interest in adapting vision-language models (VLMs) to image and third-person video classification due to their success in zero-shot recognition. However, the adaptation of these models to e…

Video ClassificationZero-Shot Learning