paper-with-me

Papers

Non-Markov Policies to Reduce Sequential Failures in Robot Bin Picking

2020-07-20 · Kate Sanders, Michael Danielczuk, Jeffrey Mahler, Ajay Tanwani, Ken Goldberg

A new generation of automated bin picking systems using deep learning is evolving to support increasing demand for e-commerce. To accommodate a wide variety of products, many automated systems include multiple gripper types and/or tool changers. However, for some objects, sequential grasp failures are common: when a computed grasp fails to lift and remove the object, the bin is often left unchanged; as the sensor input is consistent, the system retries the same grasp over and over, resulting in a significant reduction in mean successful picks per hour (MPPH). Based on an empirical study of sequential failures, we characterize a class of "sequential failure objects" (SFOs) -- objects prone to sequential failures based on a novel taxonomy. We then propose three non-Markov picking policies that incorporate memory of past failures to modify subsequent actions. Simulation experiments on SFO models and the EGAD dataset suggest that the non-Markov policies significantly outperform the Markov policy in terms of the sequential failure rate and MPPH. In physical experiments on 50 heaps of 12 SFOs the most effective Non-Markov policy increased MPPH over the Dex-Net Markov policy by 107%.

📄 PDF Abstract BibTeX arXiv:2007.10420

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Robots that Collaborate: Sequential Asymmetric Imitation for Learning Coupled Robot Policies

2026-06-15 · Yincong Chen, Ranpeng Qiu, Zihao Li, Yanan Zhou 외 arxiv

Collaborative mobile manipulation requires robots to coordinate with a partially observed partner while physically interacting through shared objects. This is difficult because failures often arise not from poor local sk…

Robot Manipulation

Reliable Robotic Task Execution in the Face of Anomalies

2025-10-27 · Bharath Santhanam, Alex Mitrevski, Santosh Thoduka, Sebastian Houben 외 arxiv

Learned robot policies have consistently been shown to be versatile, but they typically have no built-in mechanism for handling the complexity of open environments, making them prone to execution failures; this implies t…

Anomaly Detection

ELASTIC: Efficiently Learning to Adaptively Scale Test-Time Compute for Generative Control Policies

2026-06-30 · Andrew Zou Li, Gokul Swamy, Yonatan Bisk, Andrea Bajcsy arxiv

Generative control policies (GCPs), such as diffusion policies and flow-based vision-language-action models, enable test-time scaling in robot control. Test-time compute can be allocated along two axes: sequential scalin…

Reinforcement LearningRobot Manipulation

Acting in Delayed Environments with Non-Stationary Markov Policies

2021-01-28 · ICLR 2021 1 · Esther Derman, Gal Dalal, Shie Mannor

The standard Markov Decision Process (MDP) formulation hinges on the assumption that an action is executed immediately after it was chosen. However, assuming it is often unrealistic and can lead to catastrophic failures …

Cloud ComputingQ-Learning

Sequential Dexterity: Chaining Dexterous Policies for Long-Horizon Manipulation

2023-09-02 · Yuanpei Chen, Chen Wang, Li Fei-Fei, C. Karen Liu

Many real-world manipulation tasks consist of a series of subtasks that are significantly different from one another. Such long-horizon, complex tasks highlight the potential of dexterous hands, which possess adaptabilit…

Reinforcement Learning (RL)