paper-with-me

홈 › Papers

Wyvern: An Agentic Framework for Generating Grounded Multimodal Reports

2026-08-14 · Beatrice Alessandra Motetti, Emilien Guandalino, Daniele Jahier Pagliari, Alessio Burrello, Lorenz K. Müller, Konstantin Berestizshevsky, Lukas Cavigelli arxiv

In the current artificial intelligence-driven innovation era, the pace of knowledge growth is accelerating, and is hard to keep up with. While generative models are increasingly used to synthesize content, they often lack in information grounding. To address these peculiarities of our time, we propose Wyvern, a multi-agent framework for the automated generation of grounded, multimodal technical reports. Wyvern allows for the generation of multimodal outputs, integrating images, tables, and text with supporting references in a unified report. Additionally, a particular focus is placed on the grounding of the content, with the implementation of a claims auto-revision stage. We conduct a human evaluation study to assess the quality of our proposed framework. The results show that the figures' informativeness is perceived as superior to that of a recent baseline in 87% of cases. Furthermore, Wyvern's reports are rated as more useful than those produced by three alternative methods in 63% to 100% of instances. We also carry out automatic evaluations showing that Wyvern gains up to 2.3$\times$ in citation recall and 1.6$\times$ in citation precision with respect to the baselines.

📄 PDF Abstract BibTeX arXiv:2608.14446

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Unify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesis

2026-03-31 · Shuang Chen, Quanxin Shou, Hangting Chen, Yucheng Zhou 외 arxiv

Unified multimodal models provide a natural and promising architecture for understanding diverse and complex real-world knowledge while generating high-quality images. However, they still rely primarily on frozen paramet…

Image Generation

Deep-Reporter: Deep Research for Grounded Multimodal Long-Form Generation

2026-04-12 · Fangda Ye, Zhifei Xie, Yuxin Hu, Yihang Yin 외 arxiv

Recent agentic search frameworks enable deep research via iterative planning and retrieval, reducing hallucinations and enhancing factual grounding. However, they remain text-centric, overlooking the multimodal evidence …

multimodal generation

ReLook: Vision-Grounded RL with a Multimodal LLM Critic for Agentic Web Coding

2025-10-13 · Yuhang Li, Chenchen Zhang, Ruilin Lv, Ao Liu 외 arxiv

While Large Language Models (LLMs) excel at algorithmic code generation, they struggle with front-end development, where correctness is judged on rendered pixels and interaction. We present ReLook, an agentic, vision-gro…

Reinforcement LearningCode Generation

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents

2025-12-03 · Reuben Tan, Baolin Peng, Zhengyuan Yang, Hao Cheng 외 arxiv

Agentic reasoning models trained with multimodal reinforcement learning (MMRL) have become increasingly capable, yet they are almost universally optimized using sparse, outcome-based rewards computed based on the final a…

Reinforcement LearningMultimodal ReasoningSpatial Reasoning

PresentAgent-2: Towards Generalist Multimodal Presentation Agents

2026-05-12 · Wei Wu, Ziyang Xu, Zeyu Zhang, Yang Zhao 외 arxiv

Presentation generation is moving beyond static slide creation toward end-to-end presentation video generation with research grounding, multimodal media, and interactive delivery. We introduce PresentAgent-2, an agentic …

Video Generation