paper-with-me

홈 › Papers

RaMen: Multi-Strategy Multi-Modal Learning for Bundle Construction

2025-07-18 · Huy-Son Nguyen, Quang-Huy Nguyen, Duc-Hoang Pham, Duc-Trong Le, Hoang-Quynh Le, Padipat Sitkrongwong, Atsuhiro Takasu, Masoud Mansoury

Existing studies on bundle construction have relied merely on user feedback via bipartite graphs or enhanced item representations using semantic information. These approaches fail to capture elaborate relations hidden in real-world bundle structures, resulting in suboptimal bundle representations. To overcome this limitation, we propose RaMen, a novel method that provides a holistic multi-strategy approach for bundle construction. RaMen utilizes both intrinsic (characteristics) and extrinsic (collaborative signals) information to model bundle structures through Explicit Strategy-aware Learning (ESL) and Implicit Strategy-aware Learning (ISL). ESL employs task-specific attention mechanisms to encode multi-modal data and direct collaborative relations between items, thereby explicitly capturing essential bundle features. Moreover, ISL computes hyperedge dependencies and hypergraph message passing to uncover shared latent intents among groups of items. Integrating diverse strategies enables RaMen to learn more comprehensive and robust bundle representations. Meanwhile, Multi-strategy Alignment & Discrimination module is employed to facilitate knowledge transfer between learning strategies and ensure discrimination between items/bundles. Extensive experiments demonstrate the effectiveness of RaMen over state-of-the-art models on various domains, justifying valuable insights into complex item set problems.

📄 PDF Abstract BibTeX arXiv:2507.14361

Code (1)

Rec4Fun/RaMen 공식 구현

Tasks

Transfer Learning

Similar Papers 제목 키워드 기반

Leveraging Multimodal Features and Item-level User Feedback for Bundle Construction

2023-10-28 · Yunshan Ma, Xiaohao Liu, Yinwei Wei, Zhulin Tao 외

Automatic bundle construction is a crucial prerequisite step in various bundle-aware online services. Previous approaches are mostly designed to model the bundling strategy of existing bundles. However, it is hard to acq…

Contrastive LearningRepresentation Learning

RAMEN: Resolution-Adjustable Multimodal Encoder for Earth Observation

2025-12-04 · Nicolas Houdré, Diego Marcos, Hugo Riffaud de Turckheim, Dino Ienco 외 arxiv

Earth observation (EO) data spans a wide range of spatial, spectral, and temporal resolutions, from high-resolution optical imagery to low resolution multispectral products or radar time series. While recent foundation m…

Charon: a FrameNet Annotation Tool for Multimodal Corpora

2022-05-24 · LREC (LAW) 2022 6 · Frederico Belcavello, Marcelo Viridiano, Ely Edison Matos, Tiago Timponi Torrent

This paper presents Charon, a web tool for annotating multimodal corpora with FrameNet categories. Annotation can be made for corpora containing both static images and video sequences paired - or not - with text sequence…

Fine-tuning Multimodal Large Language Models for Product Bundling

2024-07-16 · Xiaohao Liu, Jie Wu, Zhulin Tao, Yunshan Ma 외

Recent advances in product bundling have leveraged multimodal information through sophisticated encoders, but remain constrained by limited semantic understanding and a narrow scope of knowledge. Therefore, some attempts…

In-Context LearningMultiple-choice

Multimodal Frame Identification with Multilingual Evaluation

2018-06-01 · NAACL 2018 6 · Teresa Botschen, Iryna Gurevych, Jan-Christoph Klie, Hatem Mousselly-Sergieh 외

An essential step in FrameNet Semantic Role Labeling is the Frame Identification (FrameId) task, which aims at disambiguating a situation around a predicate. Whilst current FrameId methods rely on textual representations…

Common Sense ReasoningSemantic Role LabelingWord Embeddings