paper-with-me

Papers

A Multimodal Fusion Model Leveraging MLP Mixer and Handcrafted Features-based Deep Learning Networks for Facial Palsy Detection

2025-03-13 · Heng Yim Nicole Oo, Min Hun Lee, Jeong Hoon Lim

Algorithmic detection of facial palsy offers the potential to improve current practices, which usually involve labor-intensive and subjective assessments by clinicians. In this paper, we present a multimodal fusion-based deep learning model that utilizes an MLP mixer-based model to process unstructured data (i.e. RGB images or images with facial line segments) and a feed-forward neural network to process structured data (i.e. facial landmark coordinates, features of facial expressions, or handcrafted features) for detecting facial palsy. We then contribute to a study to analyze the effect of different data modalities and the benefits of a multimodal fusion-based approach using videos of 20 facial palsy patients and 20 healthy subjects. Our multimodal fusion model achieved 96.00 F1, which is significantly higher than the feed-forward neural network trained on handcrafted features alone (82.80 F1) and an MLP mixer-based model trained on raw RGB images (89.00 F1).

📄 PDF Abstract BibTeX arXiv:2503.10371

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Unified Cross-Modal Attention-Mixer Based Structural-Functional Connectomics Fusion for Neuropsychiatric Disorder Diagnosis

2025-05-21 · Badhan Mazumder, Lei Wu, Vince D. Calhoun, Dong Hye Ye

Gaining insights into the structural and functional mechanisms of the brain has been a longstanding focus in neuroscience research, particularly in the context of understanding and treating neuropsychiatric disorders suc…

DiagnosticMultimodal Deep Learning

TRI-DEP: A Trimodal Comparative Study for Depression Detection Using Speech, Text, and EEG

2025-10-16 · Annisaa Fitri Nurfidausi, Eleonora Mancini, Paolo Torroni arxiv

Depression is a widespread mental health disorder, yet its automatic detection remains challenging. Prior work has explored unimodal and multimodal approaches, with multimodal systems showing promise by leveraging comple…

A Multimodal Deep Learning Framework for Edema Classification Using HCT and Clinical Data

2026-03-20 · Aram Ansary Ogholbake, Hannah Choi, Spencer Brandenburg, Alyssa Antuna 외 arxiv

We propose AttentionMixer, a unified deep learning framework for multimodal detection of brain edema that combines structural head CT (HCT) with routine clinical metadata. While HCT provides rich spatial information, cli…

Multimodal Deep Learning

MixMAS: A Framework for Sampling-Based Mixer Architecture Search for Multimodal Fusion and Learning

2024-12-24 · Abdelmadjid Chergui, Grigor Bezirganyan, Sana Sellami, Laure Berti-ÉQuille 외

Choosing a suitable deep learning architecture for multimodal data fusion is a challenging task, as it requires the effective integration and processing of diverse data types, each with distinct structures and characteri…

Benchmarking

MIXER: Multiattribute, Multiway Fusion of Uncertain Pairwise Affinities

2022-10-15 · Parker C. Lusk, Kaveh Fathian, Jonathan P. How

We present a multiway fusion algorithm capable of directly processing uncertain pairwise affinities. In contrast to existing works that require initial pairwise associations, our MIXER algorithm improves accuracy by leve…

Binarization