paper-with-me

홈 › Papers

DiffVL: Scaling Up Soft Body Manipulation using Vision-Language Driven Differentiable Physics

2023-12-11 · NeurIPS 2023 11 · Zhiao Huang, Feng Chen, Yewen Pu, Chunru Lin, Hao Su, Chuang Gan

Combining gradient-based trajectory optimization with differentiable physics simulation is an efficient technique for solving soft-body manipulation problems. Using a well-crafted optimization objective, the solver can quickly converge onto a valid trajectory. However, writing the appropriate objective functions requires expert knowledge, making it difficult to collect a large set of naturalistic problems from non-expert users. We introduce DiffVL, a method that enables non-expert users to communicate soft-body manipulation tasks -- a combination of vision and natural language, given in multiple stages -- that can be readily leveraged by a differential physics solver. We have developed GUI tools that enable non-expert users to specify 100 tasks inspired by real-life soft-body manipulations from online videos, which we'll make public. We leverage large language models to translate task descriptions into machine-interpretable optimization objectives. The optimization objectives can help differentiable physics solvers to solve these long-horizon multistage tasks that are challenging for previous baselines.

📄 PDF Abstract BibTeX arXiv:2312.06408

Code (0)

등록된 구현이 없습니다.

Tasks

valid

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

X-DiffVLA: X-Embodied Diffusion Action Heads for Vision-Language-Action Models

2026-05-24 · Boyu Li, Chaoyi Xu, Haoqi Yuan, Xinrun Xu 외 arxiv

Learning universal policies from cross-embodied data remains a fundamental challenge in robotics. Although Vision-Language-Action (VLA) models are pre-trained on large and diverse datasets, they typically rely on embodim…

DiffVLA++: Bridging Cognitive Reasoning and End-to-End Driving through Metric-Guided Alignment

2025-10-20 · Yu Gao, Anqing Jiang, Yiru Wang, Wang Jijun 외 arxiv

Conventional end-to-end (E2E) driving models are effective at generating physically plausible trajectories, but often fail to generalize to long-tail scenarios due to the lack of essential world knowledge to understand a…

Autonomous Driving

SWIM: Vision-Language-Grounded Soft Whole-Body Interactive Manipulation

2026-09-15 · Tingcong Liu, Aye Phyu Phyu Aung, Junjie Xiong, Siyi Ma 외 arxiv

Soft and continuum robots enable manipulation through distributed body deformation and contact, yet translating language and visual context into executable whole-body actuation remains a fundamental challenge. We present…

Do Rigid-Body Simulators Dream of Soft Robots? Learning Contact-Rich Manipulation for Tendon-Driven Continuum Robots

2026-06-21 · Chengnan Shentu, Nicholas Baldassini, Tongjia Zheng, Priyanka Rao 외 arxiv

Learning contact-rich, whole-body manipulation for soft continuum robots is held back by the lack of simulation infrastructure that has accelerated rigid-robot manipulation. Existing soft robot simulators are physically …

Robot Manipulation

Efficient Generation of Targeted and Transferable Adversarial Examples for Vision-Language Models Via Diffusion Models

2024-04-16 · Qi Guo, Shanmin Pang, Xiaojun Jia, Yang Liu 외

Adversarial attacks, particularly \textbf{targeted} transfer-based attacks, can be used to assess the adversarial robustness of large visual-language models (VLMs), allowing for a more thorough examination of potential s…

Adversarial DefenseAdversarial Robustness