paper-with-me

Papers

Stealing Creator's Workflow: A Creator-Inspired Agentic Framework with Iterative Feedback Loop for Improved Scientific Short-form Generation

2025-04-26 · Jong Inn Park, Maanas Taneja, Qianwen Wang, Dongyeop Kang

Generating engaging, accurate short-form videos from scientific papers is challenging due to content complexity and the gap between expert authors and readers. Existing end-to-end methods often suffer from factual inaccuracies and visual artifacts, limiting their utility for scientific dissemination. To address these issues, we propose SciTalk, a novel multi-LLM agentic framework, grounding videos in various sources, such as text, figures, visual styles, and avatars. Inspired by content creators' workflows, SciTalk uses specialized agents for content summarization, visual scene planning, and text and layout editing, and incorporates an iterative feedback mechanism where video agents simulate user roles to give feedback on generated videos from previous iterations and refine generation prompts. Experimental evaluations show that SciTalk outperforms simple prompting methods in generating scientifically accurate and engaging content over the refined loop of video generation. Although preliminary results are still not yet matching human creators' quality, our framework provides valuable insights into the challenges and benefits of feedback-driven video generation. Our code, data, and generated videos will be publicly available.

📄 PDF Abstract BibTeX arXiv:2504.18805

Code (0)

등록된 구현이 없습니다.

Tasks

FormVideo Generation

Similar Papers 제목 키워드 기반

VisionCreator: A Native Visual-Generation Agentic Model with Understanding, Thinking, Planning and Creation

2026-03-03 · Jinxiang Lai, Zexin Lu, Jiajun He, Rongwei Quan 외 arxiv

Visual content creation tasks demand a nuanced understanding of design conventions and creative workflows-capabilities challenging for general models, while workflow-based agents lack specialized knowledge for autonomous…

Reinforcement Learning

Prompt-Driven Agentic Video Editing System: Autonomous Comprehension of Long-Form, Story-Driven Media

2025-09-20 · Zihan Ding, Xinyi Wang, Junlong Chen, Per Ola Kristensson 외 arxiv

Creators struggle to edit long-form, narrative-rich videos not because of UI complexity, but due to the cognitive demands of searching, storyboarding, and sequencing hours of footage. Existing transcript- or embedding-ba…

VisionCreator-R1: A Reflection-Enhanced Native Visual-Generation Agentic Model

2026-03-09 · Jinxiang Lai, Wenzhe Zhao, Zexin Lu, Hualei Zhang 외 arxiv

Visual content generation has advanced from single-image to multi-image workflows, yet existing agents remain largely plan-driven and lack systematic reflection mechanisms to correct mid-trajectory visual errors. To addr…

Reinforcement Learning

AnimAgents: Coordinating Multi-Stage Animation Pre-Production with Human-Multi-Agent Collaboration

2025-11-22 · Wen-Fan Wang, Chien-Ting Lu, Jin Ping Ng, Yi-Ting Chiu 외 arxiv

Animation pre-production lays the foundation of an animated film by transforming initial concepts into a coherent blueprint across interdependent stages such as ideation, scripting, design, and storyboarding. While gener…

Image Generation

Navigating the Open-Source Model Ecosystem: An Empirical Study of Creator Practices in Artistic Image Generation

2026-07-12 · Yiluo Wei, Yupeng He, Qiming Ye, Gareth Tyson arxiv

The open-sourcing of powerful image generation models has created a vibrant ecosystem where creators curate and combine a vast array of community-contributed models. This practice stands in sharp contrast to using closed…

Image Generation