paper-with-me

Papers

Dynamic Cross Attention for Audio-Visual Person Verification

2024-03-07 · R. Gnana Praveen, Jahangir Alam

Although person or identity verification has been predominantly explored using individual modalities such as face and voice, audio-visual fusion has recently shown immense potential to outperform unimodal approaches. Audio and visual modalities are often expected to pose strong complementary relationships, which plays a crucial role in effective audio-visual fusion. However, they may not always strongly complement each other, they may also exhibit weak complementary relationships, resulting in poor audio-visual feature representations. In this paper, we propose a Dynamic Cross-Attention (DCA) model that can dynamically select the cross-attended or unattended features on the fly based on the strong or weak complementary relationships, respectively, across audio and visual modalities. In particular, a conditional gating layer is designed to evaluate the contribution of the cross-attention mechanism and choose cross-attended features only when they exhibit strong complementary relationships, otherwise unattended features. Extensive experiments are conducted on the Voxceleb1 dataset to demonstrate the robustness of the proposed model. Results indicate that the proposed model consistently improves the performance on multiple variants of cross-attention while outperforming the state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:2403.04661

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Audio-Visual Person Verification based on Recursive Fusion of Joint Cross-Attention

2024-03-07 · R. Gnana Praveen, Jahangir Alam

Person or identity verification has been recently gaining a lot of attention using audio-visual fusion as faces and voices share close associations with each other. Conventional approaches based on audio-visual fusion re…

PersonaTalk: Bring Attention to Your Persona in Visual Dubbing

2024-09-09 · Longhao Zhang, Shuang Liang, Zhipeng Ge, Tianshu Hu

For audio-driven visual dubbing, it remains a considerable challenge to uphold and highlight speaker's persona while synthesizing accurate lip synchronization. Existing methods fall short of capturing speaker's unique sp…

To React or not to React: End-to-End Visual Pose Forecasting for Personalized Avatar during Dyadic Conversations

2019-10-05 · Chaitanya Ahuja, Shugao Ma, Louis-Philippe Morency, Yaser Sheikh

Non verbal behaviours such as gestures, facial expressions, body posture, and para-linguistic cues have been shown to complement or clarify verbal messages. Hence to improve telepresence, in form of an avatar, it is impo…

Cross-Attention is Not Always Needed: Dynamic Cross-Attention for Audio-Visual Dimensional Emotion Recognition

2024-03-28 · R. Gnana Praveen, Jahangir Alam

In video-based emotion recognition, audio and visual modalities are often expected to have a complementary relationship, which is widely explored using cross-attention. However, they may also exhibit weak complementary r…

Emotion Recognition

Enhancing Personality Recognition by Comparing the Predictive Power of Traits, Facets, and Nuances

2026-02-05 · Amir Ansari, Jana Subirana, Bruna Silva, Sergio Escalera 외 arxiv

Personality is a complex, hierarchical construct typically assessed through item-level questionnaires aggregated into broad trait scores. Personality recognition models aim to infer personality traits from different sour…