paper-with-me

Papers

CAPE: Contrastive Action-conditioned Parallel Encoding for Embodied Planning

2026-06-05 · Cong Chen, Haowen Wang, Zhixiang Zhang, Pei Ren, Zhengping Che arxiv

Embodied agents need to predict the future consequences of candidate actions in order to plan effectively before execution. Existing visual dynamics models learn by reconstructing future visual states or rolling out dense latent representations, which spreads learning capacity across visually salient but planning-irrelevant content rather than the action-conditioned changes that drive manipulation outcomes. We propose CAPE, a Contrastive Action-conditioned Parallel Encoding framework that learns visual dynamics by distinguishing the future outcomes induced by different action sequences. Given an initial observation and a candidate action sequence, CAPE decodes the full future latent trajectory in a single forward pass and is trained with a Goal-Convergent Contrastive Objective that aligns predictions corresponding to the same future outcome while separating those corresponding to different outcomes. On real-world DROID and zero-shot transfer to RoboCasa, CAPE substantially outperforms prior baselines on future-state retrieval, offline action matching, and closed-loop planning, while notably reducing planning-time inference cost at long prediction horizons.

📄 PDF Abstract BibTeX arXiv:2606.07304

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Direction-Conditioned Policies via Compositional Subgoal Scoring for Online Goal-Conditioned Reinforcement Learning

2026-06-15 · Swaminathan S K, Damiya Gondha, Theyanesh Eswaramoorthy Rajahkrishnan, Aritra Hazra arxiv

Hamilton-Jacobi-Bellman theory implies that the optimal goal-conditioned action depends on the goal only through the gradient of the goal-reaching distance at the current state, yet standard online GCRL still conditions …

Reinforcement Learning

Structure-guided molecular design with contrastive 3D protein-ligand learning

2026-04-21 · Carles Navarro, Philipp Tholke, Gianni de Fabritiis arxiv

Structure-based drug discovery faces the dual challenge of accurately capturing 3D protein-ligand interactions while navigating ultra-large chemical spaces to identify synthetically accessible candidates. In this work, w…

Contrastive LearningDrug Discovery

Parallel Context-of-Experts Decoding for Retrieval Augmented Generation

2026-01-13 · Giulio Corallo, Paolo Papotti arxiv

Retrieval Augmented Generation faces a trade-off: concatenating documents in a long prompt enables multi-document reasoning but creates prefill bottlenecks, while encoding document KV caches separately offers speed but b…

Breaking the Reasoning Horizon in Entity Alignment Foundation Models

2026-01-29 · Yuanning Cui, Zequn Sun, Wei Hu, Kexuan Xin 외 arxiv

Entity alignment (EA) is critical for knowledge graph (KG) fusion. Existing EA models lack transferability and are incapable of aligning unseen KGs without retraining. While using graph foundation models (GFMs) offer a s…

Entity AlignmentLink Prediction

Anchor-Conditioned Compositional Control for Landscape Image Generation

2026-06-01 · Gadha Lekshmi P, Govind Arun, Rohith Syam, Ahmed Elgammal arxiv

Image generative models, though widely used as creative tools, offer limited support for the kind of compositional control that photographers and visual artists routinely exercise. This paper presents early results on an…

Image Generation