paper-with-me

홈 › Papers

Towards Verifiable Multimodal Deep Research: A Multi-Agent Harness for Interleaved Report Generation

2026-05-28 · Chenghao Zhang, Guanting Dong, Yufan Liu, Tong Zhao, Xiaoxi Li, Zhicheng Dou arxiv

Large Language Models (LLMs) have advanced autonomous agents from deep search, which retrieves concise factual answers, to deep research, which synthesizes scattered evidence into long-form reports. However, verifiable multimodal deep research remains challenging due to open-ended synthesis without deterministic ground truth and the need to interleave textual arguments with visual evidence. We propose Ptah, a multi-agent harness for interleaved report generation. Ptah orchestrates the lifecycle from user query to rendered web report through planning, research, and writing stages, where specialized agents construct visual-aware plans, collect claim-grounded evidence, maintain source-aligned images in a Visual Working Memory, and compose reports through declarative multimodal tool use. A verifier agent serves as the harness's acceptance function, enforcing factual grounding, citation fidelity, and cross-modal consistency throughout the workflow. We further introduce PtahEval, an evaluation protocol that augments existing benchmarks with image-level and presentation-level assessments. Experiments on deep research benchmarks show that Ptah produces more reliable, visually informative, and usable human-facing multimodal reports than strong baselines. Our code is released at https://github.com/SnowNation101/Ptah

📄 PDF Abstract BibTeX arXiv:2605.29861

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Code as Agent Harness

2026-05-18 · Xuying Ning, Katherine Tieu, Dongqi Fu, Tianxin Wei 외 arxiv

Recent large language models (LLMs) have demonstrated strong capabilities in understanding and generating code, from competitive programming to repository-level software engineering. In emerging agentic systems, code is …

FinanceHarness: Autonomous Financial Deep Research Framework

2026-07-30 · Yijia Xiao, Rujun Han, Yanfei Chen, Zifeng Wang 외 arxiv

Powered by advances in LLMs and autonomous agents, deep research has become one of the most widely adopted agentic products. However, most deep research systems write general-purpose reports, which are inadequate for fin…

WeAgent-MMGenEdit: A Full-Stack Recipe for Multimodal Agentic Image Generation and Editing

2026-09-04 · Hui Zhang, Zongkai Liu, Liqiang Niu, Juntao Liu 외 arxiv

Image generation and editing models have advanced rapidly, yet remain unreliable when prompts require external world knowledge. Bounded and long-tail parametric knowledge prevents direct or reason-then-generate approache…

Image GenerationImage Editing

DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces

2026-08-04 · Boyan Li, Zhuowen Liang, Yupeng Xie, Xiaotian Lin 외 arxiv

Data agents enable natural-language analytics over organizational workspaces, where relevant evidence may be scattered across databases, structured files, long documents, and multimedia. Existing benchmarks largely isola…

GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents

2026-04-08 · Mingyu Ouyang, Siyuan Hu, Kevin Qinghong Lin, Hwee Tou Ng 외 arxiv

Towards an embodied generalist for real-world interaction, Multimodal Large Language Model (MLLM) agents still suffer from challenging latency, sparse feedback, and irreversible mistakes. Video games offer an ideal testb…

Action Parsing