paper-with-me

Papers

MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome

2026-03-30 · Fangda Ye, Yuxin Hu, Pengxiang Zhu, Yibo Li, Ziqi Jin, Yao Xiao, Yibo Wang, Lei Wang, Zhen Zhang, Lu Wang, Yue Deng, Bin Wang, Yifan Zhang, Liangcai Su, Xinyu Wang, He Zhao, Chen Wei, Qiang Ren, Bryan Hooi, An Bo, Shuicheng Yan, Lidong Bing arxiv

Recent progress in deep research systems has been impressive, but evaluation still lags behind real user needs. Existing benchmarks predominantly assess final reports using fixed rubrics, failing to evaluate the underlying research process. Most also offer limited multimodal coverage, rely on synthetic tasks that do not reflect real-world query complexity, and cannot be refreshed as knowledge evolves. To address these gaps, we introduce MiroEval, a benchmark and evaluation framework for deep research systems. The benchmark comprises 100 tasks (70 text-only, 30 multimodal), all grounded in real user needs and constructed via a dual-path pipeline that supports periodic updates, enabling a live and evolving setting. The proposed evaluation suite assesses deep research systems along three complementary dimensions: adaptive synthesis quality evaluation with task-specific rubrics, agentic factuality verification via active retrieval and reasoning over both web sources and multimodal attachments, and process-centric evaluation audits how the system searches, reasons, and refines throughout its investigation. Evaluation across 13 systems yields three principal findings: the three evaluation dimensions capture complementary aspects of system capability, with each revealing distinct strengths and weaknesses across systems; process quality serves as a reliable predictor of overall outcome while revealing weaknesses invisible to output-level metrics; and multimodal tasks pose substantially greater challenges, with most systems declining by 3 to 10 points. The MiroThinker series achieves the most balanced performance, with MiroThinker-H1 ranking the highest overall in both settings. Human verification and robustness results confirm the reliability of the benchmark and evaluation framework. MiroEval provides a holistic diagnostic tool for the next generation of deep research agents.

📄 PDF Abstract BibTeX arXiv:2603.28407

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

EpiBench: Benchmarking Multi-turn Research Workflows for Multimodal Agents

2026-04-07 · Xuan Dong, Huanyang Zheng, Tianhao Niu, Zhe Han 외 arxiv

Scientific research follows multi-turn, multi-step workflows that require proactively searching the literature, consulting figures and tables, and integrating evidence across papers to align experimental settings and sup…

FormFactory: An Interactive Benchmarking Suite for Multimodal Form-Filling Agents

2025-06-02 · Bobo Li, Yuheng Wang, Hao Fei, Juncheng Li 외

Online form filling is a common yet labor-intensive task involving extensive keyboard and mouse interactions. Despite the long-standing vision of automating this process with "one click", existing tools remain largely ru…

BenchmarkingForm

VisBrowse-Bench: Benchmarking Visual-Native Search for Multimodal Browsing Agents

2026-03-17 · Zhengbo Zhang, Jinbo Su, Zhaowen Zhou, Changtao Miao 외 arxiv

The rapid advancement of Multimodal Large Language Models (MLLMs) has enabled browsing agents to acquire and reason over multimodal information in the real world. But existing benchmarks suffer from two limitations: insu…

Visual ReasoningImage Retrieval

Sketchtopia: A Dataset and Foundational Agents for Benchmarking Asynchronous Multimodal Communication with Iconic Feedback

2025-01-01 · CVPR 2025 1 · Mohd Hozaifa Khan, Ravi Kiran Sarvadevabhatla

We introduce Sketchtopia, a large-scale dataset and AI framework designed to explore goal-driven, multimodal communication through asynchronous interactions in a Pictionary-inspired setup. Sketchtopia captures natura…

Benchmarking

OpenOmni: A Collaborative Open Source Tool for Building Future-Ready Multimodal Conversational Agents

2024-08-06 · Qiang Sun, Yuanyi Luo, Sirui Li, Wenxiao Zhang 외

Multimodal conversational agents are highly desirable because they offer natural and human-like interaction. However, there is a lack of comprehensive end-to-end solutions to support collaborative development and benchma…

BenchmarkingRetrieval-augmented GenerationSpeech-to-Text