paper-with-me

홈 › Papers

Unimodal Face Classification with Multimodal Training

2021-12-08 · Wenbin Teng, Chongyang Bai

Face recognition is a crucial task in various multimedia applications such as security check, credential access and motion sensing games. However, the task is challenging when an input face is noisy (e.g. poor-condition RGB image) or lacks certain information (e.g. 3D face without color). In this work, we propose a Multimodal Training Unimodal Test (MTUT) framework for robust face classification, which exploits the cross-modality relationship during training and applies it as a complementary of the imperfect single modality input during testing. Technically, during training, the framework (1) builds both intra-modality and cross-modality autoencoders with the aid of facial attributes to learn latent embeddings as multimodal descriptors, (2) proposes a novel multimodal embedding divergence loss to align the heterogeneous features from different modalities, which also adaptively avoids the useless modality (if any) from confusing the model. This way, the learned autoencoders can generate robust embeddings in single-modality face classification on test stage. We evaluate our framework in two face classification datasets and two kinds of testing input: (1) poor-condition image and (2) point cloud or 3D face mesh, when both 2D and 3D modalities are available for training. We experimentally show that our MTUT framework consistently outperforms ten baselines on 2D and 3D settings of both datasets.

📄 PDF Abstract BibTeX arXiv:2112.04182

Code (1)

wbteng9526/mtut_fr 공식 구현 pytorch

Tasks

ClassificationFace Recognition

Similar Papers 제목 키워드 기반

CSA: Data-efficient Mapping of Unimodal Features to Multimodal Features

2024-10-10 · Po-han Li, Sandeep P. Chinchali, Ufuk Topcu

Multimodal encoders like CLIP excel in tasks such as zero-shot image classification and cross-modal retrieval. However, they require excessive training data. We propose canonical similarity analysis (CSA), which uses two…

Cross-Modal RetrievalGPUimage-classificationImage Classification+1

Enhancing Unimodal Latent Representations in Multimodal VAEs through Iterative Amortized Inference

2024-10-15 · Yuta Oshima, Masahiro Suzuki, Yutaka Matsuo

Multimodal variational autoencoders (VAEs) aim to capture shared latent representations by integrating information from different data modalities. A significant challenge is accurately inferring representations from any …

Unimodal Intermediate Training for Multimodal Meme Sentiment Classification

2023-08-01 · Muzhaffar Hazman, Susan Mckeever, Josephine Griffith

Internet Memes remain a challenging form of user-generated content for automated sentiment classification. The availability of labelled memes is a barrier to developing sentiment classifiers of multimodal memes. To addre…

ClassificationSentiment AnalysisSentiment Classification

Rethinking Vision Transformer and Masked Autoencoder in Multimodal Face Anti-Spoofing

2023-02-11 · Zitong Yu, Rizhao Cai, Yawen Cui, Xin Liu 외

Recently, vision transformer (ViT) based multimodal learning methods have been proposed to improve the robustness of face anti-spoofing (FAS) systems. However, there are still no works to explore the fundamental natures …

Face Anti-Spoofing

UniS-MMC: Multimodal Classification via Unimodality-supervised Multimodal Contrastive Learning

2023-05-16 · Heqing Zou, Meng Shen, Chen Chen, Yuchen Hu 외

Multimodal learning aims to imitate human beings to acquire complementary information from multiple modalities for various downstream tasks. However, traditional aggregation-based multimodal fusion methods ignore the int…

Contrastive LearningImage-text Classificationtext-classificationText Classification