paper-with-me

Papers

Phase-Localized Curation Does Not Help: A Negative Result on Per-Phase Metric Selection for Demonstration Filtering

2026-06-13 · Aarav Bedi arxiv

Manipulation demonstrations have temporal phase structure, and a natural hypothesis is that demonstration-curation metrics should be applied within phases rather than globally. The idea is to segment each trajectory into phases, score each phase with the metric that is locally most informative, and then aggregate. This follows directly from prior work showing that a single global metric can be the best detector of a defect and yet the worst curator of the resulting policy. We test the per-phase hypothesis on three contact-rich LIBERO pick-and-place tasks with a controlled early-release structural defect, comparing phase-gated curation against the same metrics applied uniformly and against a strong single global metric. Across all three tasks and five random seeds per condition, phase-gated curation is never the best curation strategy, and it is the worst of the three on two of the three tasks (Task 1: 86.0 vs. 92.0 for global; Task 3: 22.7 vs. 48.0 for uniform). We trace the failure to a concrete mechanism. When the defect signal is concentrated in a single phase, rank-aggregating across phases dilutes that signal with uninformative scores from defect-free phases, selecting a worse demonstration subset than simply applying the defect-informative metric everywhere. We further show that the per-phase metric selection does not transfer across tasks, since no phase shares a winning metric between any two tasks, so the selection cannot be reused and must be re-derived per task from a noisy sweep. These results bound a plausible and previously untested method, and they argue that practitioners should prefer identifying a single defect-informative metric over decomposing curation by phase. We release the full pipeline, all metric implementations, and per-seed results.

📄 PDF Abstract BibTeX arXiv:2606.15064

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RSRCC: A Remote Sensing Regional Change Comprehension Benchmark Constructed via Retrieval-Augmented Best-of-N Ranking

2026-04-22 · Roie Kazoom, Yotam Gigi, George Leifman, Tomer Shekel 외 arxiv

Traditional change detection identifies where changes occur, but does not explain what changed in natural language. Existing remote sensing change captioning datasets typically describe overall image-level differences, l…

Semantic SegmentationChange Detection

A Survey on Data Curation for Visual Contrastive Learning: Why Crafting Effective Positive and Negative Pairs Matters

2025-02-12 · Shasvat Desai, Debasmita Ghose, Deep Chakraborty

Visual contrastive learning aims to learn representations by contrasting similar (positive) and dissimilar (negative) pairs of data samples. The design of these pairs significantly impacts representation quality, trainin…

Contrastive Learning

Phoenix-VL 1.5 Medium Technical Report

2026-05-11 · Team Phoenix, :, Arka Ray, Askar Ali Mohamed Jawad 외 arxiv

We introduce Phoenix-VL 1.5 Medium, a 123B-parameter natively multimodal and multilingual foundation model, adapted to regional languages and the Singapore context. Developed as a sovereign AI asset, it demonstrates that…

Domain Adaptation

Not All Skills Help: Measuring and Repairing Agent Knowledge

2026-06-13 · Yixuan Wang, Yiyang Zhou, Yiming Liang, Congyu Zhang 외 arxiv

LLM agents can improve without weight updates by accumulating natural-language skills from experience, but current systems entrust every decision about which skills to keep and how to apply them to LLM judgment alone. We…

CaTE Data Curation for Trustworthy AI

2025-08-20 · Mary Versa Clemens-Sewall, Christopher Cervantes, Emma Rafkin, J. Neil Otte 외 arxiv

This report provides practical guidance to teams designing or developing AI-enabled systems for how to promote trustworthiness during the data curation phase of development. In this report, the authors first define data,…