paper-with-me

Papers

CommerceMM: Large-Scale Commerce MultiModal Representation Learning with Omni Retrieval

2022-02-15 · Licheng Yu, Jun Chen, Animesh Sinha, Mengjiao MJ Wang, Hugo Chen, Tamara L. Berg, Ning Zhang

We introduce CommerceMM - a multimodal model capable of providing a diverse and granular understanding of commerce topics associated to the given piece of content (image, text, image+text), and having the capability to generalize to a wide range of tasks, including Multimodal Categorization, Image-Text Retrieval, Query-to-Product Retrieval, Image-to-Product Retrieval, etc. We follow the pre-training + fine-tuning training regime and present 5 effective pre-training tasks on image-text pairs. To embrace more common and diverse commerce data with text-to-multimodal, image-to-multimodal, and multimodal-to-multimodal mapping, we propose another 9 novel cross-modal and cross-pair retrieval tasks, called Omni-Retrieval pre-training. The pre-training is conducted in an efficient manner with only two forward/backward updates for the combined 14 tasks. Extensive experiments and analysis show the effectiveness of each task. When combining all pre-training tasks, our model achieves state-of-the-art performance on 7 commerce-related downstream tasks after fine-tuning. Additionally, we propose a novel approach of modality randomization to dynamically adjust our model under different efficiency constraints.

📄 PDF Abstract BibTeX arXiv:2202.07247

Code (0)

등록된 구현이 없습니다.

Tasks

Image-text RetrievalRepresentation LearningRetrievalText Retrieval

Similar Papers 제목 키워드 기반

Captions Speak Louder than Images (CASLIE): Generalizing Foundation Models for E-commerce from High-quality Multimodal Instruction Data

2024-10-22 · Xinyi Ling, Bo Peng, Hanwen Du, Zhihui Zhu 외

Leveraging multimodal data to drive breakthroughs in e-commerce applications through Multimodal Foundation Models (MFMs) is gaining increasing attention from the research community. However, there are significant challen…

AFMRL: Attribute-Enhanced Fine-Grained Multi-Modal Representation Learning in E-commerce

2026-04-22 · Biao Zhang, Lixin Chen, Bin Zhang, Zongwei Wang 외 arxiv

Multimodal representation is crucial for E-commerce tasks such as identical product retrieval. Large representation models (e.g., VLM2Vec) demonstrate strong multimodal understanding capabilities, yet they struggle with …

Representation LearningContrastive Learning

Adapting Vision-Language Models for E-commerce Understanding at Scale

2026-02-12 · Matteo Nulli, Vladimir Orshulevich, Tala Bazazo, Christian Herold 외 arxiv

E-commerce product understanding demands by nature, strong multimodal comprehension from text, images, and structured attributes. General-purpose Vision-Language Models (VLMs) enable generalizable multimodal latent model…

Instruction FollowingAttribute Extraction

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding

2025-11-16 · Zhanheng Nie, Chenghan Fu, Daoze Zhang, Junxian Wu 외 arxiv

Recent Multimodal Large Language Models (MLLMs) have significantly advanced e-commerce product understanding. However, they still face three challenges: (i) the modality imbalance induced by modality mixed training; (ii)…

Representation Learning

MOON Embedding: Multimodal Representation Learning for E-commerce Search Advertising

2025-11-14 · Chenghan Fu, Daoze Zhang, Yukang Lin, Zhanheng Nie 외 arxiv

We introduce MOON, our comprehensive set of sustainable iterative practices for multimodal representation learning for e-commerce applications. MOON has already been fully deployed across all stages of Taobao search adve…

Representation Learning