paper-with-me

홈 › Papers

Beyond Visual Similarity: Rule-Guided Multimodal Clustering with explicit domain rules

2025-09-24 · Kishor Datta Gupta, Mohd Ariful Haque, Marufa Kamal, Ahmed Rafi Hasan, Md. Mahfuzur Rahman, Roy George arxiv

Traditional clustering techniques often rely solely on similarity in the input data, limiting their ability to capture structural or semantic constraints that are critical in many domains. We introduce the Domain Aware Rule Triggered Variational Autoencoder (DARTVAE), a rule guided multimodal clustering framework that incorporates domain specific constraints directly into the representation learning process. DARTVAE extends the VAE architecture by embedding explicit rules, semantic representations, and data driven features into a unified latent space, while enforcing constraint compliance through rule consistency and violation penalties in the loss function. Unlike conventional clustering methods that rely only on visual similarity or apply rules as post hoc filters, DARTVAE treats rules as first class learning signals. The rules are generated by LLMs, structured into knowledge graphs, and enforced through a loss function combining reconstruction, KL divergence, consistency, and violation penalties. Experiments on aircraft and automotive datasets demonstrate that rule guided clustering produces more operationally meaningful and interpretable clusters for example, isolating UAVs, unifying stealth aircraft, or separating SUVs from sedans while improving traditional clustering metrics. However, the framework faces challenges: LLM generated rules may hallucinate or conflict, excessive rules risk overfitting, and scaling to complex domains increases computational and consistency difficulties. By combining rule encodings with learned representations, DARTVAE achieves more meaningful and consistent clustering outcomes than purely data driven models, highlighting the utility of constraint guided multimodal clustering for complex, knowledge intensive settings.

📄 PDF Abstract BibTeX arXiv:2509.20501

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningKnowledge Graphs

Similar Papers 제목 키워드 기반

GRIP: Feedback-Guided Prompt Retrieval for Large Multimodal Models

2026-06-10 · Garvita Allabadi, Matteo Sodano, Roberto Estevão, Yuxiong Wang 외 arxiv

In-Context Learning (ICL) has become a powerful mechanism for adapting Large Language Models (LLMs) to new tasks without fine-tuning. Extending this concept to Large Multimodal Models (LMMs), Multimodal In-Context Learni…

Visual Question Answering

FinCriticalED: A Visual Benchmark for Financial Fact-Level OCR

2025-11-19 · Yueru He, Xueqing Peng, Yupeng Cao, Yan Wang 외 arxiv

Recent progress in multimodal large language models (MLLMs) has substantially improved document understanding, yet strong optical character recognition (OCR) performance on surface metrics does not guarantee faithful pre…

Beyond Where to Look: Trajectory-Guided Reinforcement Learning for Multimodal RLVR

2026-03-27 · Jinda Lu, Junkang Wu, Jinghan Li, Kexin Huang 외 arxiv

Recent advances in Reinforcement Learning with Verifiable Rewards (RLVR) for multimodal large language models (MLLMs) have mainly focused on improving final answer correctness and strengthening visual grounding. However,…

Reinforcement LearningMultimodal ReasoningLogical ReasoningVisual Grounding

FlashVLM: Text-Guided Visual Token Selection for Large Multimodal Models

2025-12-23 · Kaitong Cai, Jusheng Zhang, Jing Yang, Yijia Fan 외 arxiv

Large vision-language models (VLMs) typically process hundreds or thousands of visual tokens per image or video frame, incurring quadratic attention cost and substantial redundancy. Existing token reduction methods often…

VetClaw: An Edge-Cloud Multimodal Agentic System for Veterinary Disease Screening

2026-07-28 · Syed Mhamudul Hasan, Anas AlSobeh, Hussein Zangoti, Abdur R. Shahid arxiv

We present VetClaw, an edge-cloud multimodal agentic system for early veterinary disease screening. VetClaw uses a camera module as an edge sensing device and sends captured images, together with optional symptom descrip…

Image Classification