paper-with-me

홈 › Papers

Helping Figures Tell their Story! Paper-Grounded Video Generation Explaining Complex Scientific Figures

2026-06-10 · Ishani Mondal, Javad Baghirov, Jordan Boyd-Graber arxiv

Scientific figures compress complex pipelines into a single canvas, yet understanding them requires paper-grounded, step-by-step narration aligned with visual highlights a capability missing from current video generation systems and benchmarks. To address this, we introduce paper-grounded figure-to-video generation: generating narrated, region-grounded walkthrough videos from a figure and its paper. We propose MINARD (Multimodal Interpretation of Narrated Architecture via Region Decomposition), a pipeline that generates paper-grounded narrations and sequentially grounds them to figure regions. We also release FigTalk, a benchmark with new sequential and component-level grounding metrics derived. On FigTalk, MINARD generates humanlike, paper-faithful narrations and outperforms narration-conditioned figure spatial grounding compared to existing approaches in both automatic and human evaluation

📄 PDF Abstract BibTeX arXiv:2606.12576

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

The network signature of constellation line figures

2021-10-19 · Doina Bucur

In traditional astronomies across the world, groups of stars in the night sky were linked into constellations -- symbolic representations rich in meaning and with practical roles. In some sky cultures, constellations are…

Cultural Vocal Bursts Intensity Prediction

Poses of People in Art: A Data Set for Human Pose Estimation in Digital Art History

2023-01-12 · Stefanie Schneider, Ricarda Vollmer

Throughout the history of art, the pose, as the holistic abstraction of the human body's expression, has proven to be a constant in numerous studies. However, due to the enormous amount of data that so far had to be proc…

2D Human Pose Estimation2D Object DetectionMulti-Person Pose EstimationObject Detection+1

Conveying the Predicted Future to Users: A Case Study of Story Plot Prediction

2023-02-17 · Chieh-Yang Huang, Saniya Naphade, Kavya Laalasa Karanam, Ting-Hao 'Kenneth' Huang

Creative writing is hard: Novelists struggle with writer's block daily. While automatic story generation has advanced recently, it is treated as a "toy task" for advancing artificial intelligence rather than helping peop…

ARCStory ContinuationStory Generation

Inferring Implicit 3D Representations from Human Figures on Pictorial Maps

2022-08-30 · Raimund Schnürer, A. Cengiz Öztireli, Magnus Heitzler, René Sieber 외

In this work, we present an automated workflow to bring human figures, one of the most frequently appearing entities on pictorial maps, to the third dimension. Our workflow is based on training data and neural networks f…

3D ReconstructionSingle-View 3D Reconstruction

Linking Art through Human Poses

2019-07-08 · Tomas Jenicek, Ondřej Chum

We address the discovery of composition transfer in artworks based on their visual content. Automated analysis of large art collections, which are growing as a result of art digitization among museums and galleries, is a…

Content-Based Image RetrievalImage RetrievalRetrieval