paper-with-me

Papers

TSEmbed: Unlocking Task Scaling in Universal Multimodal Embeddings

2026-03-05 · Yebo Wu, Feng Liu, Ziwei Xie, Zhiyuan Liu, Changwang Zhang, Jun Wang, Li Li arxiv

Despite the exceptional reasoning capabilities of Multimodal Large Language Models (MLLMs), their adaptation into universal embedding models is significantly impeded by task conflict. To address this, we propose TSEmbed, a universal multimodal embedding framework that synergizes Mixture-of-Experts (MoE) with Low-Rank Adaptation (LoRA) to explicitly disentangle conflicting task objectives. Moreover, we introduce Expert-Aware Negative Sampling (EANS), a novel strategy that leverages expert routing distributions as an intrinsic proxy for semantic similarity. By dynamically prioritizing informative hard negatives that share expert activation patterns with the query, EANS effectively sharpens the model's discriminative power and refines embedding boundaries. To ensure training stability, we further devise a two-stage learning paradigm that solidifies expert specialization before optimizing representations via EANS. TSEmbed achieves state-of-the-art performance on both the Massive Multimodal Embedding Benchmark (MMEB) and real-world industrial production datasets, laying a foundation for task-level scaling in universal multimodal embeddings.

📄 PDF Abstract BibTeX arXiv:2603.04772

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic Similarity

Similar Papers 제목 키워드 기반

Test-Time Scaling in Multimodal Foundation Models: A Comprehensive Survey of Generation and Reasoning

2026-06-06 · Cong Wan, Ying He, Zhongzhan Huang, Hefeng Wu arxiv

Test-time Scaling (TTS) has emerged as a pivotal research direction for enhancing model performance by dynamically allocating computational resources during inference. Recent advancements have adapted this paradigm to Mu…

Multimodal Reasoning

Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models

2026-02-24 · Christian Simon, Masato Ishii, Wei-Yao Wang, Koichi Saito 외 arxiv

Scaling multimodal alignment between video and audio is challenging, particularly due to limited data and the mismatch between text descriptions and frame-level video information. In this work, we tackle the scaling chal…

Audio Generation

Generative Universal Verifier as Multimodal Meta-Reasoner

2025-10-15 · Xinchen Zhang, Xiaoying Zhang, Youbin Wu, Yanbin Cao 외 arxiv

We introduce Generative Universal Verifier, a novel concept and plugin designed for next-generation multimodal reasoning in vision-language models and unified multimodal models, providing the fundamental capability of re…

Multimodal ReasoningImage Generation

Emu3: Next-Token Prediction is All You Need

2024-09-27 · Xinlong Wang, Xiaosong Zhang, Zhengxiong Luo, Quan Sun 외

While next-token prediction is considered a promising path towards artificial general intelligence, it has struggled to excel in multimodal tasks, which are still dominated by diffusion models (e.g., Stable Diffusion) an…

AllImage GenerationPrediction+2

Unlocking Cross-Lingual Sentiment Analysis through Emoji Interpretation: A Multimodal Generative AI Approach

2024-12-23 · Rafid Ishrak Jahan, Heng Fan, Haihua Chen, Yunhe Feng

Emojis have become ubiquitous in online communication, serving as a universal medium to convey emotions and decorative elements. Their widespread use transcends language and cultural barriers, enhancing understanding and…

Sentiment Analysis