paper-with-me

홈 › Papers

MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines

2024-09-19 · Dongzhi Jiang, Renrui Zhang, Ziyu Guo, Yanmin Wu, Jiayi Lei, Pengshuo Qiu, Pan Lu, Zehui Chen, Chaoyou Fu, Guanglu Song, Peng Gao, Yu Liu, Chunyuan Li, Hongsheng Li

The advent of Large Language Models (LLMs) has paved the way for AI search engines, e.g., SearchGPT, showcasing a new paradigm in human-internet interaction. However, most current AI search engines are limited to text-only settings, neglecting the multimodal user queries and the text-image interleaved nature of website information. Recently, Large Multimodal Models (LMMs) have made impressive strides. Yet, whether they can function as AI search engines remains under-explored, leaving the potential of LMMs in multimodal search an open question. To this end, we first design a delicate pipeline, MMSearch-Engine, to empower any LMMs with multimodal search capabilities. On top of this, we introduce MMSearch, a comprehensive evaluation benchmark to assess the multimodal search performance of LMMs. The curated dataset contains 300 manually collected instances spanning 14 subfields, which involves no overlap with the current LMMs' training data, ensuring the correct answer can only be obtained within searching. By using MMSearch-Engine, the LMMs are evaluated by performing three individual tasks (requery, rerank, and summarization), and one challenging end-to-end task with a complete searching process. We conduct extensive experiments on closed-source and open-source LMMs. Among all tested models, GPT-4o with MMSearch-Engine achieves the best results, which surpasses the commercial product, Perplexity Pro, in the end-to-end task, demonstrating the effectiveness of our proposed pipeline. We further present error analysis to unveil current LMMs still struggle to fully grasp the multimodal search tasks, and conduct ablation study to indicate the potential of scaling test-time computation for AI search engine. We hope MMSearch may provide unique insights to guide the future development of multimodal AI search engine. Project Page: https://mmsearch.github.io

📄 PDF Abstract BibTeX arXiv:2409.12959

Code (0)

등록된 구현이 없습니다.

Tasks

Benchmarking

Similar Papers 제목 키워드 기반

MMSearch-Plus: Benchmarking Provenance-Aware Search for Multimodal Browsing Agents

2025-08-29 · Xijia Tao, Yihua Teng, Xinxing Su, Xinyu Fu 외 arxiv

Existing multimodal browsing benchmarks often fail to require genuine multimodal reasoning, as many tasks can be solved with text-only heuristics without vision-in-the-loop verification. We introduce MMSearch-Plus, a 311…

Multimodal ReasoningText Retrieval

ProMMSearchAgent: A Generalizable Multimodal Search Agent Trained with Process-Oriented Rewards

2026-04-22 · Wentao Yan, Shengqin Wang, Huichi Zhou, Yihang Chen 외 arxiv

Training multimodal agents via reinforcement learning for knowledge-intensive visual reasoning is fundamentally hindered by the extreme sparsity of outcome-based supervision and the unpredictability of live web environme…

Reinforcement LearningVisual Reasoning

DeepMMSearch-R1: Empowering Multimodal LLMs in Multimodal Web Search

2025-10-14 · Kartik Narayan, Yang Xu, Tian Cao, Kavya Nerella 외 arxiv

Multimodal Large Language Models (MLLMs) in real-world applications require access to external knowledge sources and must remain responsive to the dynamic and ever-changing real-world information in order to address info…

Reinforcement Learning

DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents

2026-04-21 · Shengqin Wang, Wentao Yan, Huichi Zhou, Yihang Chen 외 arxiv

Agentic multimodal models have garnered significant attention for their ability to leverage external tools to tackle complex tasks. However, it is observed that such agents often meet premature interaction collapse, caus…

WeAgent-MMSearch: Native Text-Vision Interaction for Multimodal Search Agents

2026-08-28 · Zongkai Liu, Hui Zhang, Liqiang Niu, Zhen Cao 외 arxiv

Multimodal search agents extend parametric knowledge with newly emerging and long-tail evidence from the open web. Yet many existing agentic search environments often expose retrieved evidence only as text and omit tool-…