paper-with-me

Papers

ESMC: MLLM-Based Embedding Selection for Explainable Multiple Clustering

2025-11-30 · Xinyue Wang, Yuheng Jia, Hui Liu, Junhui Hou arxiv

Typical deep clustering methods, while achieving notable progress, can only provide one clustering result per dataset. This limitation arises from their assumption of a fixed underlying data distribution, which may fail to meet user needs and provide unsatisfactory clustering outcomes. Our work investigates how multi-modal large language models (MLLMs) can be leveraged to achieve user-driven clustering, emphasizing their adaptability to user-specified semantic requirements. However, directly using MLLM output for clustering has risks for producing unstructured and generic image descriptions instead of feature-specific and concrete ones. To address these issues, our method first discovers that MLLMs' hidden states of text tokens are strongly related to the corresponding features, and leverages these embeddings to perform clusterings from any user-defined criteria. We also employ a lightweight clustering head augmented with pseudo-label learning, significantly enhancing clustering accuracy. Extensive experiments demonstrate its competitive performance on diverse datasets and metrics.

📄 PDF Abstract BibTeX arXiv:2512.00725

Code (0)

등록된 구현이 없습니다.

Tasks

Deep Clustering

Similar Papers 제목 키워드 기반

A Generative Physics-Informed Reinforcement Learning-Based Approach for Construction of Representative Drive Cycle

2025-06-09 · Amirreza Yasami, Mohammadali Tofigh, Mahdi Shahbakhti, Charles Robert Koch

Accurate driving cycle construction is crucial for vehicle design, fuel economy analysis, and environmental impact assessments. A generative Physics-Informed Expected SARSA-Monte Carlo (PIESMC) approach that constructs r…

Creative4U: MLLMs-based Advertising Creative Image Selector with Comparative Reasoning

2025-08-18 · Yukang Lin, Xiang Zhang, Shichang Jia, Bowen Wan 외 arxiv

Creative image in advertising is the heart and soul of e-commerce platform. An eye-catching creative image can enhance the shopping experience for users, boosting income for advertisers and advertising revenue for platfo…

Reinforcement Learning

FusionAgent: A Multimodal Agent with Dynamic Model Selection for Human Recognition

2026-03-27 · Jie Zhu, Xiao Guo, Yiyang Su, Anil Jain 외 arxiv

Model fusion is a key strategy for robust recognition in unconstrained scenarios, as different models provide complementary strengths. This is especially important for whole-body human recognition, where biometric cues s…

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering

2026-03-19 · Dan Ben-Ami, Gabriele Serussi, Kobi Cohen, Chaim Baskin arxiv

Long-form video question answering requires reasoning over extended temporal contexts, making frame selection a critical bottleneck for multi-modal large language models (MLLMs) bound by finite context windows. Within th…

Video Question AnsweringSpeech Recognition

ESMC: Entire Space Multi-Task Model for Post-Click Conversion Rate via Parameter Constraint

2023-07-18 · Zhenhao Jiang, Biao Zeng, Hao Feng, Jin Liu 외

Large-scale online recommender system spreads all over the Internet being in charge of two basic tasks: Click-Through Rate (CTR) and Post-Click Conversion Rate (CVR) estimations. However, traditional CVR estimators suffe…

Decision MakingRecommendation SystemsSelection bias