paper-with-me

홈 › Papers

Early Warning Signals for OpenVLA Failure under Visual Distribution Shift

2026-06-29 · Dipesh Tharu Mahato, Rachel Ren arxiv

Vision Language Action models combine perception, language grounding, and control in a single policy, but their failures are hard to diagnose once visual conditions shift. We test whether OpenVLA feedforward activations contain linearly decodable information about near term task failure in LIBERO manipulation rollouts. The policy is fixed throughout. We log internal activations during execution and fit lightweight monitors after the rollouts are collected. Occlusion is the main controlled stress test. It reduces OpenVLA success from $57\%$ to $17\%$ over $100$ episodes per condition. Under this shift, a logistic probe at layer 16 reaches AUROC $0.972$ and AUPRC $0.352$ for predicting failure within a $15$ step horizon. It outperforms both a mean difference direction and an action disagreement baseline. A sparse layer sweep finds uneven decodability across depth: layer 16 is strongest among the tested layers, layer 8 remains informative, and layer 10 is weaker. To check whether the monitor is just an occlusion detector, we also evaluate color shift and camera jitter without refitting. Color shift produces no failures in this setting, so it is a benign control rather than a failure benchmark. Camera jitter does induce failures, and the occlusion trained monitor remains above random. The result is deliberately limited: OpenVLA internal states contain failure relevant structure under controlled perceptual shift, but these experiments do not establish a causal mechanism, task held out generalization, or a deployable recovery system.

📄 PDF Abstract BibTeX arXiv:2606.29699

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Early Diagnosis of Wasted Computation in Multi-Agent LLM Systems via Failure-Aware Observability

2026-05-31 · Xianyou Li, Weiran Yan, Yichao Wu, Penghao Liang 외 arxiv

Failure-aware observability diagnoses wasted computation in multi-agent LLM systems before final-answer evaluation can explain what went wrong. We propose a trace-based framework for a three-agent architecture -- orchest…

TRACER: Early Failure Detection for Task-Oriented Dialogue

2026-07-04 · Erfan Nourbakhsh, Rocky Slavin, Ke Yang, Anthony Rios arxiv

Task-oriented dialogue systems often fail before the final breakdown is obvious, but most evaluation only measures failure after the conversation has already gone wrong. We present TRACER, a method for early failure dete…

Task-Oriented Dialogue Systems

When GPUs Fail Quietly: Observability-Aware Early Warning Beyond Numeric Telemetry

2026-03-17 · Michael Bidollahkhani, Freja Nordsiek, Julian M. Kunkel arxiv

GPU nodes are central to modern HPC and AI workloads, yet many failures do not manifest as immediate hard faults. While some instabilities emerge gradually as weak thermal or efficiency drift, a significant class occurs …

Learning from the past: predicting critical transitions with machine learning trained on surrogates of historical data

2024-10-13 · Zhiqin Ma, Chunhua Zeng, Yi-Cheng Zhang, Thomas M. Bury

Complex systems can undergo critical transitions, where slowly changing environmental conditions trigger a sudden shift to a new, potentially catastrophic state. Early warning signals for these events are crucial for dec…

Decision MakingSociologySpecificity

Early Warning Signals Appear Long Before Dropping Out: An Idiographic Approach Grounded in Complex Dynamic Systems Theory

2026-01-16 · Mohammed Saqr, Sonsoles López-Pernas, Santtu Tikka, Markus Wolfgang Hermann Spitzer arxiv

The ability to sustain engagement and recover from setbacks (i.e., resilience) -- is fundamental for learning. When resilience weakens, students are at risk of disengagement and may drop out and miss on opportunities. Th…