paper-with-me

홈 › Papers

Token-Region Guided Cross-Attention Fusion for Multimodal Affect Interpretation

2026-07-26 · Musa Tur Farazi, Nufayer Jahan Reza arxiv

Automated analysis of multimodal content on social networks has become a critical task for understanding public sentiment and information diffusion in the digital age. However, classifying internet memes remains computationally challenging due to the intricate interplay between visual cues and embedded, often stylized, text, particularly in low-resource languages like Bengali Language. This paper addresses the detection of political intent in Bengali memes by introducing Multimodal Cross-Attention Fusion framework. We first leverage a Vision-Language Model to extract high-fidelity OCR text from noisy meme images. Subsequently, we encode visual and textual features and synthesize them through a cross-modal multi-head attention mechanism that aligns semantic tokens with visual regions. We also investigate the integration of a domain-specific political lexicon as a knowledge prior. Experimental evaluation on the PoliMemeDecode1 dataset shows that our attention-based fusion significantly outperforms unimodal baselines and standard concatenation methods, achieving a state-of-the-art Macro-F1 of approximately 0.94. Interpretability analyzes further confirm that the model effectively learns to ground textual semantics in visual evidence.

📄 PDF Abstract BibTeX arXiv:2607.23493

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Mask-Guided Attention Regulation for Anatomically Consistent Counterfactual CXR Synthesis

2026-03-04 · Zichun Zhang, Weizhi Nie, Honglin Guo, Yuting Su arxiv

Counterfactual generation for chest X-rays (CXR) aims to simulate plausible pathological changes while preserving patient-specific anatomy. However, diffusion-based editing methods often suffer from structural drift, whe…

Data Augmentation

Focus-dLLM: Accelerating Long-Context Diffusion LLM Inference via Confidence-Guided Context Focusing

2026-02-02 · Lingkun Long, Yushi Huang, Shihao Bai, Ruihao Gong 외 arxiv

Diffusion Large Language Models (dLLMs) deliver strong long-context processing capability in a non-autoregressive decoding paradigm. However, the considerable computational cost of bidirectional full attention limits the…

Token Painter: Training-Free Text-Guided Image Inpainting via Mask Autoregressive Models

2025-09-28 · Longtao Jiang, Jie Huang, Mingfei Han, Lei Chen 외 arxiv

Text-guided image inpainting aims to inpaint masked image regions based on a textual prompt while preserving the background. Although diffusion-based methods have become dominant, their property of modeling the entire im…

Image Inpainting

LocInv: Localization-aware Inversion for Text-Guided Image Editing

2024-05-02 · Chuanming Tang, Kai Wang, Fei Yang, Joost Van de Weijer

Large-scale Text-to-Image (T2I) diffusion models demonstrate significant generation capabilities based on textual prompts. Based on the T2I diffusion models, text-guided image editing research aims to empower users to ma…

Denoisingtext-guided-image-editing

BindEdit: Taming Attention Leakage for Precise Multi-Object Image Editing

2026-06-17 · Chaewon Park, Soyoon Lee, Naeun Lee, Minjung Shin 외 arxiv

Real image editing enables precise manipulation of visual content, yet existing methods often fail in complex multi-object scenarios, causing semantic blending, object duplication, or incomplete edits. We attribute these…

Image Editing