paper-with-me

Papers

Generative Timelines for Instructed Visual Assembly

2024-11-19 · Alejandro Pardo, Jui-Hsien Wang, Bernard Ghanem, Josef Sivic, Bryan Russell, Fabian Caba Heilbron

The objective of this work is to manipulate visual timelines (e.g. a video) through natural language instructions, making complex timeline editing tasks accessible to non-expert or potentially even disabled users. We call this task Instructed visual assembly. This task is challenging as it requires (i) identifying relevant visual content in the input timeline as well as retrieving relevant visual content in a given input (video) collection, (ii) understanding the input natural language instruction, and (iii) performing the desired edits of the input visual timeline to produce an output timeline. To address these challenges, we propose the Timeline Assembler, a generative model trained to perform instructed visual assembly tasks. The contributions of this work are three-fold. First, we develop a large multimodal language model, which is designed to process visual content, compactly represent timelines and accurately interpret timeline editing instructions. Second, we introduce a novel method for automatically generating datasets for visual assembly tasks, enabling efficient training of our model without the need for human-labeled data. Third, we validate our approach by creating two novel datasets for image and video assembly, demonstrating that the Timeline Assembler substantially outperforms established baseline models, including the recent GPT-4o, in accurately executing complex assembly instructions across various real-world inspired scenarios.

📄 PDF Abstract BibTeX arXiv:2411.12293

Code (0)

등록된 구현이 없습니다.

Tasks

Language Modelling

Similar Papers 제목 키워드 기반

Automatic Verbal Depiction of a Brick Assembly for a Robot Instructing Humans

2022-09-01 · SIGDIAL (ACL) 2022 9 · Rami Younes, Gérard Bailly, Frederic Elisei, Damien Pellier

Verbal and nonverbal communication skills are essential for human-robot interaction, in particular when the agents are involved in a shared task. We address the specific situation when the robot is the only agent knowing…

Building LEGO Using Deep Generative Models of Graphs

2020-12-21 · Rylee Thompson, Elahe Ghalebi, Terrance DeVries, Graham W. Taylor

Generative models are now used to create a variety of high-quality digital artifacts. Yet their use in designing physical objects has received far less attention. In this paper, we advocate for the construction toy, LEGO…

TAILOR: Teaching with Active and Incremental Learning for Object Registration

2022-05-24 · Qianli Xu, Nicolas Gauthier, Wenyu Liang, Fen Fang 외

When deploying a robot to a new task, one often has to train it to detect novel objects, which is time-consuming and labor-intensive. We present TAILOR -- a method and system for object registration with active and incre…

Incremental LearningObject

Blox-Net: Generative Design-for-Robot-Assembly Using VLM Supervision, Physics Simulation, and a Robot with Reset

2024-09-25 · Andrew Goldberg, Kavish Kondap, Tianshuang Qiu, Zehan Ma 외

Generative AI systems have shown impressive capabilities in creating text, code, and images. Inspired by the rich history of research in industrial ''Design for Assembly'', we introduce a novel problem: Generative Design…

Motion Planning

Shape of Thought: Progressive Object Assembly via Visual Chain-of-Thought

2026-01-28 · Yu Huo, Siyu Zhang, Kun Zeng, Haoyue Liu 외 arxiv

Multimodal models for text-to-image generation have achieved strong visual fidelity, yet they remain brittle under compositional structural constraints, notably generative numeracy, attribute binding, and part-level rela…

Text-to-Image Generation