paper-with-me

Papers

Fuse after Align: Improving Face-Voice Association Learning via Multimodal Encoder

2024-04-15 · Chong Peng, Liqiang He, Dan Su

Today, there have been many achievements in learning the association between voice and face. However, most previous work models rely on cosine similarity or L2 distance to evaluate the likeness of voices and faces following contrastive learning, subsequently applied to retrieval and matching tasks. This method only considers the embeddings as high-dimensional vectors, utilizing a minimal scope of available information. This paper introduces a novel framework within an unsupervised setting for learning voice-face associations. By employing a multimodal encoder after contrastive learning and addressing the problem through binary classification, we can learn the implicit information within the embeddings in a more effective and varied manner. Furthermore, by introducing an effective pair selection method, we enhance the learning outcomes of both contrastive learning and the matching task. Empirical evidence demonstrates that our framework achieves state-of-the-art results in voice-face matching, verification, and retrieval tasks, improving verification by approximately 3%, matching by about 2.5%, and retrieval by around 1.3%.

📄 PDF Abstract BibTeX arXiv:2404.09509

Code (0)

등록된 구현이 없습니다.

Tasks

Binary ClassificationContrastive LearningRetrieval

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association

2025-05-22 · Abdul Hannan, Muhammad Arslan Manzoor, Shah Nawaz, Muhammad Irzam Liaqat 외

We study the task of learning association between faces and voices, which is gaining interest in the multimodal community lately. These methods suffer from the deliberate crafting of negative mining procedures as well as…

RFOP: Rethinking Fusion and Orthogonal Projection for Face-Voice Association

2025-12-02 · Abdul Hannan, Furqan Malik, Hina Jabbar, Syed Suleman Sadiq 외 arxiv

Face-voice association in multilingual environment challenge 2026 aims to investigate the face-voice association task in multilingual scenario. The challenge introduces English-German face-voice pairs to be utilized in t…

Learning Branched Fusion and Orthogonal Projection for Face-Voice Association

2022-08-22 · Muhammad Saad Saeed, Shah Nawaz, Muhammad Haris Khan, Sajid Javed 외

Recent years have seen an increased interest in establishing association between faces and voices of celebrities leveraging audio-visual information from YouTube. Prior works adopt metric learning methods to learn an emb…

Metric Learning

FaVoA: Face-Voice Association Favours Ambiguous Speaker Detection

2021-09-01 · Hugo Carneiro, Cornelius Weber, Stefan Wermter

The strong relation between face and voice can aid active speaker detection systems when faces are visible, even in difficult settings, when the face of a speaker is not clear or when there are several people in the same…

Active Speaker Detection

Fusion and Orthogonal Projection for Improved Face-Voice Association

2021-12-20 · Muhammad Saad Saeed, Muhammad Haris Khan, Shah Nawaz, Muhammad Haroon Yousaf 외

We study the problem of learning association between face and voice, which is gaining interest in the computer vision community lately. Prior works adopt pairwise or triplet loss formulations to learn an embedding space …

Cross-Modal RetrievalTriplet