paper-with-me

홈 › Papers

From Observation to Action: Latent Action-based Primitive Segmentation for VLA Pre-training in Industrial Settings

2025-11-26 · Jiajie Zhang, Sören Schwertfeger, Alexander Kleiner arxiv

We present a novel unsupervised framework to unlock vast unlabeled human demonstration data from continuous industrial video streams for Vision-Language-Action (VLA) model pre-training. Our method first trains a lightweight motion tokenizer to encode motion dynamics, then employs an unsupervised action segmenter leveraging a novel "Latent Action Energy" metric to discover and segment semantically coherent action primitives. The pipeline outputs both segmented video clips and their corresponding latent action sequences, providing structured data directly suitable for VLA pre-training. Evaluations on public benchmarks and a proprietary electric motor assembly dataset demonstrate effective segmentation of key tasks performed by humans at workstations. Further clustering and quantitative assessment via a Vision-Language Model confirm the semantic coherence of the discovered action primitives. To our knowledge, this is the first fully automated end-to-end system for extracting and organizing VLA pre-training data from unstructured industrial videos, offering a scalable solution for embodied AI integration in manufacturing.

📄 PDF Abstract BibTeX arXiv:2511.21428

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Latent Actions from Factorized Transition Effects under Agent Ambiguity

2026-06-29 · Heejeong Nam, Chandradithya S Jonnalagadda, Harshit Aggarwal, Eric Xu 외 arxiv

Latent Action Models (LAMs) learn action-like proxies from observation transitions. However, in multi-object or distractor-rich scenes, these visual effects mix agent motion with distractors, camera dynamics, and backgro…

LAC: Latent Action Composition for Skeleton-based Action Segmentation

2023-08-28 · Di Yang, Yaohui Wang, Antitza Dantcheva, Quan Kong 외

Skeleton-based action segmentation requires recognizing composable actions in untrimmed videos. Current approaches decouple this problem by first extracting local visual features from skeleton sequences and then processi…

Action SegmentationContrastive LearningSegmentationSkeleton Based Action Segmentation+1

LAC - Latent Action Composition for Skeleton-based Action Segmentation

2023-01-01 · ICCV 2023 1 · Di Yang, Yaohui Wang, Antitza Dantcheva, Quan Kong 외

Skeleton-based action segmentation requires recognizing composable actions in untrimmed videos. Current approaches decouple this problem by first extracting local visual features from skeleton sequences and then proc…

Action SegmentationContrastive LearningSegmentationSkeleton Based Action Segmentation+1

Agent Primitives: Reusable Latent Building Blocks for Multi-Agent Systems

2026-02-03 · Haibo Jin, Peng Kuang, Ye Yu, Xiaopeng Yuan 외 arxiv

While existing multi-agent systems (MAS) can handle complex problems by enabling collaboration among multiple agents, they are often highly task-specific, relying on manually crafted agent roles and interaction prompts, …

iSeg3D: An Interactive 3D Shape Segmentation Tool

2021-12-24 · Sucheng Qian, Liu Liu, Wenqiang Xu, Cewu Lu

A large-scale dataset is essential for learning good features in 3D shape understanding, but there are only a few datasets that can satisfy deep learning training. One of the major reasons is that current tools for annot…

Segmentation