paper-with-me

홈 › Papers

SIEVE: Structure-Aware Data Selection for Imitation Learning with VLA Models

2026-07-07 · Changti Wu, Bin Yu, Zhaolong Shen, Shijie Lian, Xiaopeng Lin, Cong Huang, Zhirui Zhang, Lei Zhang, Kai Chen arxiv

Vision-Language-Action (VLA) models are typically trained by imitation learning on large-scale robot demonstration datasets, but more data does not necessarily yield better policies due to redundancy, noise, and uneven coverage. Existing data selection methods often assess demonstrations at either the trajectory or state-action level, missing the reusable structures that compose long-horizon behaviors. In this paper, we propose SIEVE, a structure-aware data selection method for VLA imitation learning. SIEVE views demonstrations as compositions of reusable primitives and transition interfaces. It first discovers visuo-motor primitives from segmented trajectories, then allocates selection budgets to composition patterns by maximizing reuse-aware structural exposure under diminishing returns. Finally, it selects medoid trajectories within each composition-pattern bucket to retain central, stable, and imitation-friendly demonstrations. Experiments across multiple datasets, benchmarks, and VLA models show that SIEVE consistently outperforms competitive data selection baselines. Notably, SIEVE can surpass full-data training while using only 50% of demonstrations and 50% of training steps, suggesting that reusable structure, captured through primitives and transitions, is an important signal for efficient VLA imitation learning.

📄 PDF Abstract BibTeX arXiv:2607.06442

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LogSieve: Task-Aware CI Log Reduction for Sustainable LLM-Based Analysis

2026-01-28 · Marcus Emmanuel Barnes, Taher A. Ghaleb, Safwat Hassan arxiv

Logs are essential for understanding Continuous Integration (CI) behavior, particularly for diagnosing build failures and performance regressions. Yet their growing volume and verbosity make both manual inspection and au…

Anomaly Detection

OopsieVerse: A Safety Benchmark with Damage-Aware Simulation for Robot Manipulation

2026-06-30 · Arnav Balaji, Arpit Bahety, Sriniket Ambatipudi, Daniel Lam 외 arxiv

While robotic manipulation capabilities have advanced rapidly, physical safety remains a major barrier to deploying household robots: task success is insufficient if the robot damages itself or its surroundings. Simulati…

Reinforcement LearningRobot Manipulation

SCION: Size-aware Policy Orchestration for Nonstationary Object Caches (Long Paper Version)

2026-03-27 · Qizhi Wang arxiv

Object caches underpin cloud and edge services, but production workloads are heterogeneous, nonstationary, and throughput-constrained. Recent simple non-ML policies such as SIEVE and S3-FIFO set a strong baseline, so any…

TabSieve: Explicit In-Table Evidence Selection for Tabular Prediction

2026-02-12 · Yongyao Wang, Ziqi Miao, Lu Yang, Haonan Jia 외 arxiv

Tabular prediction can benefit from in-table rows as few-shot evidence, yet existing tabular models typically perform instance-wise inference and LLM-based prompting is often brittle. Models do not consistently leverage …

Reinforcement Learning

MetaSieve: Faster Relational Deep Learning through SQL-Based Metapath Selection

2026-08-26 · Fahim Shahriar Khan, Ashraf Aboulnaga arxiv

Relational Deep Learning (RDL) is an effective approach to machine learning over multi-table relational databases. In RDL, a database is modeled as a graph in which each row is a node and each foreign-key relation is an …

Graph Neural Network