paper-with-me

Papers

Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents

2026-05-09 · Ying Chen, Lihuang Fang, Rui Jiang, Mingxu Wang, Zhifeng Gu, Lei Yi, Jie Chen arxiv

Standard embodied evaluations do not independently score whether an agent correctly commits to task completion at episode closure, a capacity we call terminal commitment. Behaviorally distinct failures--never completing the task, completing it but failing to stop, and reporting success without sufficient evidence--collapse into the same benchmark failure. We introduce VIGIL, an evaluation framework that makes terminal commitment independently measurable. Under VIGIL's default protocol, agents observe only egocentric RGB, receive no action-success signals, and must end each episode with a semantic report checked deterministically against hidden world state. This yields two separate scores: world-state completion (W) and benchmark success (B), where B additionally requires a correct terminal report. This decoupling makes four outcome categories distinguishable: missed execution, post-attainment drift, unsupported commitment, and verified success. Across 20 models on 1,000 frozen episodes, systems with comparable W differ by up to 19.7 pp in B: one model converts achieved states into correct reports, while another with near-identical execution drifts past the goal without closing. An action-feedback intervention further tests the separation: execution-oriented signals improve W broadly, yet commitment failures persist in models that do not already ground terminal reports in the achieved state. VIGIL provides a protocol that makes terminal commitment independently visible and scorable.

📄 PDF Abstract BibTeX arXiv:2605.08747

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Learning Disentangling and Fusing Networks for Face Completion Under Structured Occlusions

2017-12-13 · Zhihang Li, Yibo Hu, Ran He

Face completion aims to generate semantically new pixels for missing facial components. It is a challenging generative task due to large variations of face appearance. This paper studies generative face completion under …

DecoderFacial InpaintingGenerative Adversarial Network

Self-Supervised Feature Learning from Partial Point Clouds via Pose Disentanglement

2022-01-09 · Meng-Shiun Tsai, Pei-Ze Chiang, Yi-Hsuan Tsai, Wei-Chen Chiu

Self-supervised learning on point clouds has gained a lot of attention recently, since it addresses the label-efficiency and domain-gap problems on point cloud tasks. In this paper, we propose a novel self-supervised fra…

DisentanglementRepresentation LearningSelf-Supervised Learning

RealDiff: Real-world 3D Shape Completion using Self-Supervised Diffusion Models

2024-09-16 · Başak Melis Öcal, Maxim Tatarchenko, Sezer Karaoglu, Theo Gevers

Point cloud completion aims to recover the complete 3D shape of an object from partial observations. While approaches relying on synthetic shape priors achieved promising results in this domain, their applicability and g…

ObjectPoint Cloud Completion

U-DuDoNet: Unpaired dual-domain network for CT metal artifact reduction

2021-03-08 · Yuanyuan Lyu, Jiajun Fu, Cheng Peng, S. Kevin Zhou

Recently, both supervised and unsupervised deep learning methods have been widely applied on the CT metal artifact reduction (MAR) task. Supervised methods such as Dual Domain Network (Du-DoNet) work well on simulation d…

DisentanglementMetal Artifact Reduction

Disentangling Structure and Aesthetics for Style-Aware Image Completion

2018-06-01 · CVPR 2018 6 · Andrew Gilbert, John Collomosse, Hailin Jin, Brian Price

Content-aware image completion or in-painting is a fundamental tool for the correction of defects or removal of objects in images. We propose a non-parametric in-painting algorithm that enforces both structural and aest…