paper-with-me

홈 › Papers

Quality-Aware Multimodal Fusion Reveals Implicit Identity in Valence-Arousal Features

2026-07-23 · Jisu Kim, Benjamin S. Riggan arxiv

Conventional face recognition relies on static appearance cues and degrades in unconstrained settings with expression variation, occlusion, and poor lighting. We hypothesize that audiovisual expression dynamics carry identity-discriminative information complementary to static appearance, and that extracting this signal requires multimodal representations robust to the variable input quality of in-the-wild video. To learn such representations, we cast multimodal valence-arousal (VA) estimation as a pretext task and propose Quality-Aware Adaptive Fusion (QAAF), which estimates per-sample, per-modality reliability and adapts each modality's contribution through learned soft gating and a quality-dependent dropout. For the problem of VA estimation, QAAF achieves an average Concordance Correlation Coefficient (CCC) of 0.472 via late fusion ensembling on Aff-wild2, improving over a baseline ensemble under the same setting (0.415) as well as a single-backbone baseline (0.288). Furthermore, the proposed QAAF demonstrates greater resilience to unavailable modalities, with only a 7.5-34.4% relative decrease in CCC when one modality is missing. We then probe whether these VA-trained features encode identity without identity-specific training. On AFEW-VA (67 actors) and YTF (1,595 subjects), VA-trained backbone features rank first among evaluated soft biometric methods, and score-level fusion with ArcFace lowers EER on both datasets (0.022 to 0.021 on AFEW-VA, 0.106 to 0.104 on YTF), correcting 68.2% of ArcFace's false accepts on AFEW-VA. These findings establish multimodal VA estimation as a soft biometric modality complementary to conventional face recognition.

📄 PDF Abstract BibTeX arXiv:2607.21347

Code (0)

등록된 구현이 없습니다.

Tasks

Face Recognition

Similar Papers 제목 키워드 기반

IMG: Calibrating Diffusion Models via Implicit Multimodal Guidance

2025-09-30 · Jiayi Guo, Chuanhao Yan, Xingqian Xu, Yulin Wang 외 arxiv

Ensuring precise multimodal alignment between diffusion-generated images and input prompts has been a long-standing challenge. Earlier works finetune diffusion weight using high-quality preference data, which tends to be…

TMTE: Effective Multimodal Graph Learning with Task-aware Modality and Topology Co-evolution

2026-03-29 · Yinlin Zhu, Xunkai Li, Di Wu, Wang Luo 외 arxiv

Multimodal-attributed graphs (MAGs) are a fundamental data structure for multimodal graph learning (MGL), enabling both graph-centric and modality-centric tasks. However, our empirical analysis reveals inherent topology …

Metric LearningGraph Learning

Provable Dynamic Fusion for Low-Quality Multimodal Data

2023-06-03 · Qingyang Zhang, Haitao Wu, Changqing Zhang, QinGhua Hu 외

The inherent challenge of multimodal fusion is to precisely capture the cross-modal correlation and flexibly conduct cross-modal interaction. To fully release the value of each modality and mitigate the influence of low-…

Quality-Aware Multimodal Biometric Recognition

2021-12-10 · Sobhan Soleymani, Ali Dabouei, Fariborz Taherkhani, Seyed Mehdi Iranmanesh 외

We present a quality-aware multimodal recognition framework that combines representations from multiple biometric traits with varying quality and number of samples to achieve increased recognition accuracy by extracting …

QASA: Quality-Aware Semantic Augmentation for Robust Multimodal Sentiment Analysis

2026-01-11 · Jiazhang Liang, Jianheng Dai, Miaosen Luo, Menghua Jiang 외 arxiv

Multimodal large language models have demonstrated strong ability in capturing semantic representations for multimodal sentiment analysis. Their capacity to learn stable and generalizable multimodal features is limited, …

Multimodal Sentiment AnalysisData Augmentation