paper-with-me

홈 › Papers

Transformer-Driven Triple Fusion Framework for Enhanced Multimodal Author Intent Classification in Low-Resource Bangla

2025-11-28 · Ariful Islam, Tanvir Mahmud, Md Rifat Hossen arxiv

The expansion of the Internet and social networks has led to an explosion of user-generated content. Author intent understanding plays a crucial role in interpreting social media content. This paper addresses author intent classification in Bangla social media posts by leveraging both textual and visual data. Recognizing limitations in previous unimodal approaches, we systematically benchmark transformer-based language models (mBERT, DistilBERT, XLM-RoBERTa) and vision architectures (ViT, Swin, SwiftFormer, ResNet, DenseNet, MobileNet), utilizing the Uddessho dataset of 3,048 posts spanning six practical intent categories. We introduce a novel intermediate fusion strategy that significantly outperforms early and late fusion on this task. Experimental results show that intermediate fusion, particularly with mBERT and Swin Transformer, achieves 84.11% macro-F1 score, establishing a new state-of-the-art with an 8.4 percentage-point improvement over prior Bangla multimodal approaches. Our analysis demonstrates that integrating visual context substantially enhances intent classification. Cross-modal feature integration at intermediate levels provides optimal balance between modality-specific representation and cross-modal learning. This research establishes new benchmarks and methodological standards for Bangla and other low-resource languages. We call our proposed framework BangACMM (Bangla Author Content MultiModal).

📄 PDF Abstract BibTeX arXiv:2511.23287

Code (0)

등록된 구현이 없습니다.

Tasks

Intent Classification

Similar Papers 제목 키워드 기반

OneHOI: Unifying Human-Object Interaction Generation and Editing

2026-04-15 · Jiun Tian Hoe, Weipeng Hu, Xudong Jiang, Yap-Peng Tan 외 arxiv

Human-Object Interaction (HOI) modelling captures how humans act upon and relate to objects, typically expressed as <person, action, object> triplets. Existing approaches split into two disjoint families: HOI generation …

Vanast: Virtual Try-On with Human Image Animation via Synthetic Triplet Supervision

2026-04-06 · Hyunsoo Cha, Wonjung Woo, Byungjun Kim, Hanbyul Joo arxiv

We present Vanast, a unified framework that generates garment-transferred human animation videos directly from a single human image, garment images, and a pose guidance video. Conventional two-stage pipelines treat image…

Virtual Try-on

A Relation-Attentive 3D Matrix Framework for Relational Triple Extraction

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Extracting relational triples from unstructured text is crucial for information extraction. Recent methods achieve considerable performance, but due to the insufficient consideration of triple global information, there i…

DecoderRelationvalid

One Pass for All: A Discrete Diffusion Model for Knowledge Graph Triple Set Prediction

2026-04-20 · Jihong Guan, Jiaqi Wang, Wengen Li, Hanchen Yang 외 arxiv

Knowledge Graphs (KGs) are composed of triples, and the goal of Knowledge Graph Completion (KGC) is to infer the missing factual triples. Traditional KGC tasks predict missing elements in a triple given one or two of its…

Knowledge Graph CompletionGraph GenerationKnowledge Graphs

TrioPose: Native Triple-Stream Diffusion Transformers for Pose-Guided Text-to-Image Generation

2026-06-05 · Dian Gu, Zhengyi Yang arxiv

Pose-guided text-to-image generation often suffers from limb distortions and feature crosstalk in complex multi-person scenarios. While existing UNet-based adapters struggle with long-range spatial dependencies, emerging…

Text-to-Image Generation