paper-with-me

홈 › Papers

A Context-aware Attention and Graph Neural Network-based Multimodal Framework for Misogyny Detection

2025-08-07 · Mohammad Zia Ur Rehman, Sufyaan Zahoor, Areeb Manzoor, Musharaf Maqbool, Nagendra Kumar arxiv

A substantial portion of offensive content on social media is directed towards women. Since the approaches for general offensive content detection face a challenge in detecting misogynistic content, it requires solutions tailored to address offensive content against women. To this end, we propose a novel multimodal framework for the detection of misogynistic and sexist content. The framework comprises three modules: the Multimodal Attention module (MANM), the Graph-based Feature Reconstruction Module (GFRM), and the Content-specific Features Learning Module (CFLM). The MANM employs adaptive gating-based multimodal context-aware attention, enabling the model to focus on relevant visual and textual information and generating contextually relevant features. The GFRM module utilizes graphs to refine features within individual modalities, while the CFLM focuses on learning text and image-specific features such as toxicity features and caption features. Additionally, we curate a set of misogynous lexicons to compute the misogyny-specific lexicon score from the text. We apply test-time augmentation in feature space to better generalize the predictions on diverse inputs. The performance of the proposed approach has been evaluated on two multimodal datasets, MAMI and MMHS150K, with 11,000 and 13,494 samples, respectively. The proposed method demonstrates an average improvement of 10.17% and 8.88% in macro-F1 over existing methods on the MAMI and MMHS150K datasets, respectively.

📄 PDF Abstract BibTeX arXiv:2508.09175

Code (0)

등록된 구현이 없습니다.

Tasks

Graph Neural Network

Similar Papers 제목 키워드 기반

Leveraging Foundation Models for Multimodal Graph-Based Action Recognition

2025-05-21 · Fatemeh Ziaeetabar, Florentin Wörgötter

Foundation models have ushered in a new era for multimodal video understanding by enabling the extraction of rich spatiotemporal and semantic representations. In this work, we introduce a novel graph-based framework that…

Action RecognitionGraph AttentionVideo Understanding

Hypergraph and Latent ODE Learning for Multimodal Root Cause Localization in Microservices

2026-05-01 · Xin Liu, Yuhang He, Sichen Zhao, Kejian Tong 외 arxiv

Root cause localization in cloud native microservice systems requires modeling complex service dependencies, irregular temporal dynamics, and heterogeneous observability data. We present HyperODE RCA, a unified framework…

AGSP-DSA: An Adaptive Graph Signal Processing Framework for Robust Multimodal Fusion with Dynamic Semantic Alignment

2026-01-26 · KV Karthikeya, Ashok Kumar Das, Shantanu Pal, Vivekananda Bhat K 외 arxiv

In this paper, we introduce an Adaptive Graph Signal Processing with Dynamic Semantic Alignment (AGSP DSA) framework to perform robust multimodal data fusion over heterogeneous sources, including text, audio, and images.…

Sentiment Analysis

Multimodal LLM Integrated Semantic Communications for 6G Immersive Experiences

2025-07-07 · Yusong Zhang, Yuxuan Sun, Lei Guo, Wei Chen 외 arxiv

6G networks promise revolutionary immersive communication experiences including augmented reality (AR), virtual reality (VR), and holographic communications. These applications demand high-dimensional multimodal data tra…

Visual Question AnsweringImage Generation

Geometry-Aware Superpixel Graph Transformer with Metadata for Skin Lesion Classification

2026-06-18 · Muhammad Azeem, Tanveer Hussain, Amr Ahmed, Ardhendu Behera arxiv

Automated skin cancer classification from dermoscopic images remains challenging due to heterogeneous lesion structure, strong intra-class variability, and subtle visual differences between benign and malignant cases. Ex…

Skin Lesion ClassificationSkin Cancer ClassificationMultimodal ReasoningGraph Learning