paper-with-me

홈 › Papers

Adversarial Stress Testing of SPARK Humanoid Safety Filters

2026-05-18 · Saurav Ghosh, Abdou Sow, Luke Zhang arxiv

Humanoid robots are difficult to deploy safely because they have high-dimensional bodies, many collision constraints, and must operate near people and obstacles. Safety filters help by modifying a nominal control action when it may violate collision-avoidance constraints. Still, nominal benchmark scores do not fully show how these filters behave in harder environments. In this work, we study the robustness of SPARK humanoid safety filters through replication and stress testing. We replicate the SPARK benchmark case G1SportMode_D1_WG_SO_v1 in MuJoCo and evaluate RSSA, RSSS, SSA, CBF, PFM, and SMA under controlled random seeds. We also built a post-processing pipeline that converts raw SPARK logs into goal-tracking, minimum-distance, and collision-step metrics. Our results show that some methods track the goal more closely, while others reduce collision steps more effectively. The stress tests further indicate that safety behavior can change under obstacle crowding, noisy distance estimates, and delayed obstacle information. These findings suggest that humanoid autonomy should be evaluated beyond nominal performance, using metrics that expose failure modes before deployment.

📄 PDF Abstract BibTeX arXiv:2605.19009

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Adversarial Moral Stress Testing of Large Language Models

2026-04-01 · Saeid Jamshidi, Foutse Khomh, Arghavan Moradi Dakhel, Amin Nikanjam 외 arxiv

Evaluating the ethical robustness of large language models (LLMs) deployed in software systems remains challenging, particularly under sustained adversarial user interaction. Existing safety benchmarks typically rely on …

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation

2026-08-04 · Saqib Shouqi, Abdullah Nazly, Januki Wanniarachchi, Ravisha De Alwis arxiv

Role-Playing Language Agents (RPLAs) are increasingly deployed in high-stakes applications such as healthcare assistance, customer support, and education, where maintaining consistent personas, ethical constraints, and b…

CRASH: Challenging Reinforcement-Learning Based Adversarial Scenarios For Safety Hardening

2024-11-26 · Amar Kulkarni, Shangtong Zhang, Madhur Behl

Ensuring the safety of autonomous vehicles (AVs) requires identifying rare but critical failure cases that on-road testing alone cannot discover. High-fidelity simulations provide a scalable alternative, but automaticall…

Autonomous VehiclesDeep Reinforcement Learningreinforcement-learningReinforcement Learning

Adversarial Stress Testing of Lifetime Distributions

2020-03-27 · Nozer Singpurwalla

In this paper we put forward the viewpoint that the notion of stress testing financial institutions and engineered systems can also be made viable appropos the stress testing an individual's strength of conviction in a p…

Stress-testing medical large language models reveals latent safety pathology beyond benchmark accuracy

2026-06-06 · Yuan Shen, Xiaojun Wu, Linghua Yu arxiv

Large language models (LLMs) are entering clinical practice based on benchmark accuracy that may fail to detect safety-relevant failure modes. Here we present AI-MASLD, a stress-audit framework that adapts the logic of m…

Information Extraction