paper-with-me

홈 › Papers

Pre-Manipulation Alignment Prediction with Parallel Deep State-Space and Transformer Models

2025-09-17 · Motonari Kambara, Komei Sugiura arxiv

In this work, we address the problem of predicting the future success of open-vocabulary object manipulation tasks. Conventional approaches typically determine success or failure after the action has been carried out. However, they make it difficult to prevent potential hazards and rely on failures to trigger replanning, thereby reducing the efficiency of object manipulation sequences. To overcome these challenges, we propose a model, which predicts the alignment between a pre-manipulation egocentric image with the planned trajectory and a given natural language instruction. We introduce a Multi-Level Trajectory Fusion module, which employs a state-of-the-art deep state-space model and a transformer encoder in parallel to capture multi-level time-series self-correlation within the end effector trajectory. Our experimental results indicate that the proposed method outperformed existing methods, including foundation models.

📄 PDF Abstract BibTeX arXiv:2509.13839

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ABot-M0.5: Unified Mobility-and-Manipulation World Action Model

2026-07-01 · Ronghan Chen, Yandan Yang, Zuojin Tang, Dongjie Huo 외 hf

Mobile manipulation is a key capability for general-purpose robots, yet remains challenging for current embodied learning methods. VLA policies are typically reactive and lack explicit world modeling, while existing Worl…

OASIS: Observation-Action Space Alignment via SE(3) Trajectory Prediction for Robotic Manipulation

2026-05-25 · Xinzhe Chen, Sihua Ren, Liqi Huang, Haowen Sun 외 arxiv

Recent vision-language-action (VLA) models and world action models (WAMs) advance robotic manipulation by enriching intermediate representations with auxiliary spatial features or future visual-state prediction. However,…

Trajectory Prediction

Parareal with a Learned Coarse Model for Robotic Manipulation

2019-12-12 · Wisdom Agboh, Oliver Grainger, Daniel Ruprecht, Mehmet Dogar

A key component of many robotics model-based planning and control algorithms is physics predictions, that is, forecasting a sequence of states given an initial state and a sequence of controls. This process is slow and a…

MuJoCo

BridgeVLA: Input-Output Alignment for Efficient 3D Manipulation Learning with Vision-Language Models

2025-06-09 · Peiyan Li, Yixiang Chen, Hongtao Wu, Xiao Ma 외

Recently, leveraging pre-trained vision-language models (VLMs) for building vision-language-action (VLA) models has emerged as a promising approach to effective robot manipulation learning. However, only few methods inco…

Robot ManipulationVision-Language-Action

Multi-Modal Decentralized Reinforcement Learning for Modular Reconfigurable Lunar Robots

2025-10-23 · Ashutosh Mishra, Shreya Santra, Elian Neppel, Edoardo M. Rossi Lombardi 외 arxiv

Modular reconfigurable robots suit task-specific space operations, but the combinatorial growth of morphologies hinders unified control. We propose a decentralized reinforcement learning (Dec-RL) scheme where each module…

Zero-shot GeneralizationReinforcement Learning