paper-with-me

홈 › Papers

Imagine How To Change: Explicit Procedure Modeling for Change Captioning

2026-03-06 · Jiayang Sun, Zixin Guo, Min Cao, Guibo Zhu, Jorma Laaksonen arxiv

Change captioning generates descriptions that explicitly describe the differences between two visually similar images. Existing methods operate on static image pairs, thus ignoring the rich temporal dynamics of the change procedure, which is the key to understand not only what has changed but also how it occurs. We introduce ProCap, a novel framework that reformulates change modeling from static image comparison to dynamic procedure modeling. ProCap features a two-stage design: The first stage trains a procedure encoder to learn the change procedure from a sparse set of keyframes. These keyframes are obtained by automatically generating intermediate frames to make the implicit procedural dynamics explicit and then sampling them to mitigate redundancy. Then the encoder learns to capture the latent dynamics of these keyframes via a caption-conditioned, masked reconstruction task. The second stage integrates this trained encoder within an encoder-decoder model for captioning. Instead of relying on explicit frames from the previous stage -- a process incurring computational overhead and sensitivity to visual noise -- we introduce learnable procedure queries to prompt the encoder for inferring the latent procedure representation, which the decoder then translates into text. The entire model is then trained end-to-end with a captioning loss, ensuring the encoder's output is both temporally coherent and captioning-aligned. Experiments on three datasets demonstrate the effectiveness of ProCap. Code and pre-trained models are available at https://github.com/BlueberryOreo/ProCap

📄 PDF Abstract BibTeX arXiv:2603.05969

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SCHEMA: State CHangEs MAtter for Procedure Planning in Instructional Videos

2024-03-03 · Yulei Niu, Wenliang Guo, Long Chen, Xudong Lin 외

We study the problem of procedure planning in instructional videos, which aims to make a goal-oriented sequence of action steps given partial visual state observations. The motivation of this problem is to learn a struct…

Contrastive Learning

What Changed and What Could Have Changed? State-Change Counterfactuals for Procedure-Aware Video Representation Learning

2025-03-27 · Chi-Hsi Kung, Frangil Ramirez, Juhyung Ha, Yi-Ting Chen 외

Understanding a procedural activity requires modeling both how action steps transform the scene, and how evolving scene transformations can influence the sequence of action steps, even those that are accidental or errone…

Action SegmentationcounterfactualCounterfactual ReasoningRepresentation Learning+1

A Bayesian Model of Diachronic Meaning Change

2016-01-01 · TACL 2016 1 · Lea Frermann, Mirella Lapata

Word meanings change over time and an automated procedure for extracting this information from text would be useful for historical exploratory studies, information retrieval or question answering. We present a dynamic Ba…

Change DetectionGeneral ClassificationInformation Retrievalmodel+2

Imagination Helps Visual Reasoning, But Not Yet in Latent Space

2026-02-26 · You Li, Chi Chen, Yanghao Li, Fanhu Zeng 외 arxiv

Latent visual reasoning aims to mimic human's imagination process by meditating through hidden states of Multimodal Large Language Models. While recognized as a promising paradigm for visual reasoning, the underlying mec…

Visual Reasoning

CoSIm: Commonsense Reasoning for Counterfactual Scene Imagination

2022-07-08 · NAACL 2022 7 · Hyounghun Kim, Abhay Zala, Mohit Bansal

As humans, we can modify our assumptions about a scene by imagining alternative objects or concepts in our minds. For example, we can easily anticipate the implications of the sun being overcast by rain clouds (e.g., the…

counterfactual