paper-with-me

Papers

Textual-to-Visual Iterative Self-Verification for Slide Generation

2025-02-21 · Yunqing Xu, Xinbei Ma, Jiyang Qiu, Hai Zhao

Generating presentation slides is a time-consuming task that urgently requires automation. Due to their limited flexibility and lack of automated refinement mechanisms, existing autonomous LLM-based agents face constraints in real-world applicability. We decompose the task of generating missing presentation slides into two key components: content generation and layout generation, aligning with the typical process of creating academic slides. First, we introduce a content generation approach that enhances coherence and relevance by incorporating context from surrounding slides and leveraging section retrieval strategies. For layout generation, we propose a textual-to-visual self-verification process using a LLM-based Reviewer + Refiner workflow, transforming complex textual layouts into intuitive visual formats. This modality transformation simplifies the task, enabling accurate and human-like review and refinement. Experiments show that our approach significantly outperforms baseline methods in terms of alignment, logical flow, visual appeal, and readability.

📄 PDF Abstract BibTeX arXiv:2502.15412

Code (0)

등록된 구현이 없습니다.

Tasks

Layout Generation

Similar Papers 제목 키워드 기반

AutoPresent: Designing Structured Visuals from Scratch

2025-01-01 · CVPR 2025 1 · Jiaxin Ge, Zora Zhiruo Wang, Xuhui Zhou, Yi-Hao Peng 외

Designing structured visuals such as presentation slides is essential for communicative needs, necessitating both content creation and visual planning skills. In this work, we tackle the challenge of automated slide gene…

Image Generation

Reflect to Inform: Boosting Multimodal Reasoning via Information-Gain-Driven Verification

2026-03-27 · Shuai Lv, Chang Liu, Feng Tang, Yujie Yuan 외 arxiv

Multimodal Large Language Models (MLLMs) achieve strong multimodal reasoning performance, yet we identify a recurring failure mode in long-form generation: as outputs grow longer, models progressively drift away from ima…

Multimodal Reasoning

Bridging Modality Disconnect in Self-Reflection via Closed-Loop Visually Grounded Verification

2026-02-21 · Haoyu Zhang, Yuwei Wu, Pengxiang Li, Xintong Zhang 외 arxiv

In the era of Vision-Language Models (VLMs), enhancing multimodal reasoning capabilities remains a critical challenge, particularly in handling ambiguous or complex visual inputs, where initial inferences often lead to h…

Multimodal Reasoning

Pathology-knowledge Enhanced Multi-instance Prompt Learning for Few-shot Whole Slide Image Classification

2024-07-15 · Linhao Qu, Dingkang Yang, Dan Huang, Qinhao Guo 외

Current multi-instance learning algorithms for pathology image analysis often require a substantial number of Whole Slide Images for effective training but exhibit suboptimal performance in scenarios with limited learnin…

image-classificationImage ClassificationPrompt Learningwhole slide images

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs

2025-07-22 · Yangshu Yuan, Heng Chen, Xinyi Jiang, Christian Ng 외 arxiv

The rapid advancement of Large Language Models (LLMs) and Large Vision-Language Models (LVLMs) has enhanced our ability to process and generate human language and visual information. However, these models often struggle …

Instruction FollowingMultimodal ReasoningResponse GenerationLogical Reasoning