paper-with-me

홈 › Papers

Stitch-a-Recipe: Video Demonstration from Multistep Descriptions

2025-03-18 · Chi Hsuan Wu, Kumar Ashutosh, Kristen Grauman

When obtaining visual illustrations from text descriptions, today's methods take a description with-a single text context caption, or an action description-and retrieve or generate the matching visual context. However, prior work does not permit visual illustration of multistep descriptions, e.g. a cooking recipe composed of multiple steps. Furthermore, simply handling each step description in isolation would result in an incoherent demonstration. We propose Stitch-a-Recipe, a novel retrieval-based method to assemble a video demonstration from a multistep description. The resulting video contains clips, possibly from different sources, that accurately reflect all the step descriptions, while being visually coherent. We formulate a training pipeline that creates large-scale weakly supervised data containing diverse and novel recipes and injects hard negatives that promote both correctness and coherence. Validated on in-the-wild instructional videos, Stitch-a-Recipe achieves state-of-the-art performance, with quantitative gains up to 24% as well as dramatic wins in a human preference study.

📄 PDF Abstract BibTeX arXiv:2503.13821

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Multistep Quasimetric Learning for Scalable Goal-conditioned Reinforcement Learning

2025-11-11 · Bill Chunyuan Zheng, Vivek Myers, Benjamin Eysenbach, Sergey Levine arxiv

Learning how to reach goals in an environment is a longstanding challenge in AI, yet reasoning over long horizons remains a challenge for modern methods. The key question is how to estimate the temporal distance between …

Reinforcement Learning

Procedural Text Generation from an Execution Video

2017-11-01 · IJCNLP 2017 11 · Atsushi Ushiku, Hayato Hashimoto, Atsushi Hashimoto, Shinsuke Mori

In recent years, there has been a surge of interest in automatically describing images or videos in a natural language. These descriptions are useful for image/video search, etc. In this paper, we focus on procedure exec…

Object RecognitionText GenerationVideo Captioning

Dynamic Multistep Reasoning based on Video Scene Graph for Video Question Answering

2022-07-01 · NAACL 2022 7 · Jianguo Mao, Wenbin Jiang, Xiangdong Wang, Zhifan Feng 외

Existing video question answering (video QA) models lack the capacity for deep video understanding and flexible multistep reasoning. We propose for video QA a novel model which performs dynamic multistep reasoning betwee…

Question AnsweringVideo Question AnsweringVideo Understanding

Eliminating Warping Shakes for Unsupervised Online Video Stitching

2024-03-11 · Lang Nie, Chunyu Lin, Kang Liao, Yun Zhang 외

In this paper, we retarget video stitching to an emerging issue, named warping shake, when extending image stitching to video stitching. It unveils the temporal instability of warped content in non-overlapping regions, d…

Image StitchingVideo Stabilization

StabStitch++: Unsupervised Online Video Stitching with Spatiotemporal Bidirectional Warps

2025-05-08 · Lang Nie, Chunyu Lin, Kang Liao, Yun Zhang 외

We retarget video stitching to an emerging issue, named warping shake, which unveils the temporal content shakes induced by sequentially unsmooth warps when extending image stitching to video stitching. Even if the input…

Image StitchingVideo Stabilization