paper-with-me

Papers

Instruction-based Image Manipulation by Watching How Things Move

2024-12-16 · CVPR 2025 1 · Mingdeng Cao, Xuaner Zhang, Yinqiang Zheng, Zhihao Xia

This paper introduces a novel dataset construction pipeline that samples pairs of frames from videos and uses multimodal large language models (MLLMs) to generate editing instructions for training instruction-based image manipulation models. Video frames inherently preserve the identity of subjects and scenes, ensuring consistent content preservation during editing. Additionally, video data captures diverse, natural dynamics-such as non-rigid subject motion and complex camera movements-that are difficult to model otherwise, making it an ideal source for scalable dataset construction. Using this approach, we create a new dataset to train InstructMove, a model capable of instruction-based complex manipulations that are difficult to achieve with synthetically generated datasets. Our model demonstrates state-of-the-art performance in tasks such as adjusting subject poses, rearranging elements, and altering camera perspectives.

📄 PDF Abstract BibTeX arXiv:2412.12087

Code (0)

등록된 구현이 없습니다.

Tasks

Image Manipulation

Similar Papers 제목 키워드 기반

InstructMove: A Text-Indispensable Benchmark for Instruction-Following Manipulation

2026-08-24 · Mengao Zhao, Ziang Li, Chaodong Huang, Mengchen Ma 외 arxiv

Vision-language-action (VLA) models have made general-purpose robot manipulation increasingly plausible by conditioning robot actions on natural-language instructions. A key test of such generality is whether policies ac…

Instruction FollowingRobot ManipulationSpatial Reasoning

Learning Object Manipulation Skills via Approximate State Estimation from Real Videos

2020-11-13 · Vladimír Petrík, Makarand Tapaswi, Ivan Laptev, Josef Sivic

Humans are adept at learning new tasks by watching a few instructional videos. On the other hand, robots that learn new actions either require a lot of effort through trial and error, or use expert demonstrations that ar…

ObjectState Estimation

Point and Instruct: Enabling Precise Image Editing by Unifying Direct Manipulation and Text Instructions

2024-02-05 · Alec Helbling, Seongmin Lee, Polo Chau

Machine learning has enabled the development of powerful systems capable of editing images from natural language instructions. However, in many common scenarios it is difficult for users to specify precise image transfor…

Image Manipulation

CNeuroMod-THINGS, a densely-sampled fMRI dataset for visual neuroscience

2025-07-11 · Marie St-Laurent, Basile Pinsard, Oliver Contier, Elizabeth DuPre 외 arxiv

Data-hungry neuro-AI modelling requires ever larger neuroimaging datasets. CNeuroMod-THINGS meets this need by capturing neural representations for a wide set of semantic concepts using well-characterized images in a new…

Can You Move These Over There? An LLM-based VR Mover for Supporting Object Manipulation

2025-02-04 · Xiangzhi Eric Wang, Zackary P. T. Sin, Ye Jia, Daniel Archer 외

In our daily lives, we can naturally convey instructions for the spatial manipulation of objects using words and gestures. Transposing this form of interaction into virtual reality (VR) object manipulation can be benefic…

Object