paper-with-me

홈 › Papers

Talk2Move: Reinforcement Learning for Text-Instructed Object-Level Geometric Transformation in Scenes

2026-01-05 · Jing Tan, Zhaoyang Zhang, Yantao Shen, Jiarui Cai, Shuo Yang, Jiajun Wu, Wei Xia, Zhuowen Tu, Stefano Soatto arxiv

We introduce Talk2Move, a reinforcement learning (RL) based diffusion framework for text-instructed spatial transformation of objects within scenes. Spatially manipulating objects in a scene through natural language poses a challenge for multimodal generation systems. While existing text-based manipulation methods can adjust appearance or style, they struggle to perform object-level geometric transformations-such as translating, rotating, or resizing objects-due to scarce paired supervision and pixel-level optimization limits. Talk2Move employs Group Relative Policy Optimization (GRPO) to explore geometric actions through diverse rollouts generated from input images and lightweight textual variations, removing the need for costly paired data. A spatial reward guided model aligns geometric transformations with linguistic description, while off-policy step evaluation and active step sampling improve learning efficiency by focusing on informative transformation stages. Furthermore, we design object-centric spatial rewards that evaluate displacement, rotation, and scaling behaviors directly, enabling interpretable and coherent transformations. Experiments on curated benchmarks demonstrate that Talk2Move achieves precise, consistent, and semantically faithful object transformations, outperforming existing text-guided editing approaches in both spatial accuracy and scene coherence.

📄 PDF Abstract BibTeX arXiv:2601.02356

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learningmultimodal generation

Similar Papers 제목 키워드 기반

Multi-Objective Instruction-Aware Representation Learning in Procedural Content Generation RL

2025-08-08 · Sung-Hyun Kim, Geum-Hwan Hwang, In-Chang Baek, Seo-Young Lee 외 arxiv

Recent advancements in generative modeling emphasize the importance of natural language as a highly expressive and accessible modality for controlling content generation. However, existing instructed reinforcement learni…

Multi-Label ClassificationRepresentation LearningReinforcement Learning

GAIA: Zero-shot Talking Avatar Generation

2023-11-26 · Tianyu He, Junliang Guo, Runyi Yu, Yuchi Wang 외

Zero-shot talking avatar generation aims at synthesizing natural talking videos from speech and a single portrait image. Previous methods have relied on domain-specific heuristics such as warping-based motion representat…

Diversity

Fine-tuning Transformers with Additional Context to Classify Discursive Moves in Mathematics Classrooms

2022-07-01 · NAACL (BEA) 2022 7 · Abhijit Suresh, Jennifer Jacobs, Margaret Perkoff, James H. Martin 외

“Talk moves” are specific discursive strategies used by teachers and students to facilitate conversations in which students share their thinking, and actively consider the ideas of others, and engage in rich discussions.…

Re-HOLD: Video Hand Object Interaction Reenactment via adaptive Layout-instructed Diffusion Model

2025-03-21 · CVPR 2025 1 · Yingying Fan, Quanwei Yang, Kaisiyuan Wang, Hang Zhou 외

Current digital human studies focusing on lip-syncing and body movement are no longer sufficient to meet the growing industrial demand, while human video generation techniques that support interacting with real-world env…

DisentanglementHuman-Object Interaction DetectionObjectVideo Generation

Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking Heads

2025-07-31 · Yingjie Zhou, Jiezhang Cao, Zicheng Zhang, Farong Wen 외 arxiv

Speech-driven methods for portraits are figuratively known as "Talkers" because of their capability to synthesize speaking mouth shapes and facial movements. Especially with the rapid development of the Text-to-Image (T2…