paper-with-me

홈 › Papers

Face-Voice Association with Inductive Bias for Maximum Class Separation

2026-01-20 · Marta Moscati, Oleksandr Kats, Mubashir Noman, Muhammad Zaigham Zaheer, Yufang Hou, Markus Schedl, Shah Nawaz arxiv

Face-voice association is widely studied in multimodal learning and is approached representing faces and voices with embeddings that are close for a same person and well separated from those of others. Previous work achieved this with loss functions. Recent advancements in classification have shown that the discriminative ability of embeddings can be strengthened by imposing maximum class separation as inductive bias. This technique has never been used in the domain of face-voice association, and this work aims at filling this gap. More specifically, we develop a method for face-voice association that imposes maximum class separation among multimodal representations of different speakers as an inductive bias. Through quantitative experiments we demonstrate the effectiveness of our approach, showing that it achieves SOTA performance on two task formulation of face-voice association. Furthermore, we carry out an ablation study to show that imposing inductive bias is most effective when combined with losses for inter-class orthogonality. To the best of our knowledge, this work is the first that applies and demonstrates the effectiveness of maximum class separation as an inductive bias in multimodal learning; it hence paves the way to establish a new paradigm.

📄 PDF Abstract BibTeX arXiv:2601.13651

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RFOP: Rethinking Fusion and Orthogonal Projection for Face-Voice Association

2025-12-02 · Abdul Hannan, Furqan Malik, Hina Jabbar, Syed Suleman Sadiq 외 arxiv

Face-voice association in multilingual environment challenge 2026 aims to investigate the face-voice association task in multilingual scenario. The challenge introduces English-German face-voice pairs to be utilized in t…

FaVoA: Face-Voice Association Favours Ambiguous Speaker Detection

2021-09-01 · Hugo Carneiro, Cornelius Weber, Stefan Wermter

The strong relation between face and voice can aid active speaker detection systems when faces are visible, even in difficult settings, when the face of a speaker is not clear or when there are several people in the same…

Active Speaker Detection

Face-voice Association in Multilingual Environments (FAME) Challenge 2024 Evaluation Plan

2024-04-14 · Muhammad Saad Saeed, Shah Nawaz, Muhammad Salman Tahir, Rohan Kumar Das 외

The advancements of technology have led to the use of multimodal systems in various real-world applications. Among them, the audio-visual systems are one of the widely used multimodal systems. In the recent years, associ…

Face-voice Association in Multilingual Environments (FAME) 2026 Challenge Evaluation Plan

2025-08-06 · Marta Moscati, Ahmed Abdullah, Muhammad Saad Saeed, Shah Nawaz 외 arxiv

The advancements of technology have led to the use of multimodal systems in various real-world applications. Among them, audio-visual systems are among the most widely used multimodal systems. In the recent years, associ…

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association

2025-05-22 · Abdul Hannan, Muhammad Arslan Manzoor, Shah Nawaz, Muhammad Irzam Liaqat 외

We study the task of learning association between faces and voices, which is gaining interest in the multimodal community lately. These methods suffer from the deliberate crafting of negative mining procedures as well as…