paper-with-me

Papers

SlideGen: Collaborative Multimodal Agents for Scientific Slide Generation

2025-12-04 · Xin Liang, Xiang Zhang, Yiwei Xu, Siqi Sun, Chenyu You arxiv

Generating academic slides from scientific papers is a challenging multimodal reasoning task that requires both long context understanding and deliberate visual planning. Existing approaches largely reduce it to text only summarization, overlooking the visual component and design intensive nature of slide creation. In this paper we introduce SlideGen, an agentic, modular, and visual in the loop framework for scientific paper to slide generation. SlideGen orchestrates a group of vision language agents that reason collaboratively over the document structure and semantics, producing editable PPTX slides with logical flow and compelling visual presentation. By integrating coordinated outlining, mapping, arrangement, note synthesis, and iterative refinement, our system consistently delivers slides of expert level quality. Across diverse benchmarks and strong baselines, SlideGen outperforms existing methods in visual quality, content faithfulness, and readability, positioning it as the new state of the art in automated slide generation. Our work establishes a foundation for design aware multimodal slide generation, demonstrating how agentic collaboration can bridge understanding and presentation in complex multimodal reasoning tasks.

📄 PDF Abstract BibTeX arXiv:2512.04529

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal Reasoning

Similar Papers 제목 키워드 기반

Human-Agent Collaborative Paper-to-Page Crafting

2025-10-22 · Qianli Ma, Siyu Wang, Yilin Chen, Yinhao Tang 외 arxiv

In the quest for scientific progress, communicating research is as vital as the discovery itself. Yet, researchers are often sidetracked by the manual, repetitive chore of building project webpages to make their dense pa…

VideoAgent: Personalized Synthesis of Scientific Videos

2025-09-14 · Xiao Liang, Bangxin Li, Zixuan Chen, Hanyue Zheng 외 arxiv

The technical complexity of research papers often limits their reach, necessitating more accessible formats like scientific videos to disseminate key insights through engaging narration. However, existing automated metho…

SlideBot: A Multi-Agent Framework for Generating Informative, Reliable, Multi-Modal Presentations

2025-11-12 · Eric Xie, Danielle Waterfield, Michael Kennedy, Aidong Zhang arxiv

Large Language Models (LLMs) have shown immense potential in education, automating tasks like quiz generation and content summarization. However, generating effective presentation slides introduces unique challenges due …

Code Generation

WSI-Agents: A Collaborative Multi-Agent System for Multi-Modal Whole Slide Image Analysis

2025-07-19 · Xinheng Lyu, Yuci Liang, Wenting Chen, Meidan Ding 외 arxiv

Whole slide images (WSIs) are vital in digital pathology, enabling gigapixel tissue analysis across various pathological tasks. While recent advancements in multi-modal large language models (MLLMs) allow multi-task WSI …

Exploring the Potential of Multimodal LLM with Knowledge-Intensive Multimodal ASR

2024-06-16 · Minghan Wang, Yuxia Wang, Thuy-Trang Vu, Ehsan Shareghi 외

Recent advancements in multimodal large language models (MLLMs) have made significant progress in integrating information across various modalities, yet real-world applications in educational and scientific domains remai…