paper-with-me

홈 › Papers

RoboAct-CLIP: Video-Driven Pre-training of Atomic Action Understanding for Robotics

2025-04-02 · Zhiyuan Zhang, Yuxin He, Yong Sun, Junyu Shi, Lijiang Liu, Qiang Nie

Visual Language Models (VLMs) have emerged as pivotal tools for robotic systems, enabling cross-task generalization, dynamic environmental interaction, and long-horizon planning through multimodal perception and semantic reasoning. However, existing open-source VLMs predominantly trained for generic vision-language alignment tasks fail to model temporally correlated action semantics that are crucial for robotic manipulation effectively. While current image-based fine-tuning methods partially adapt VLMs to robotic applications, they fundamentally disregard temporal evolution patterns in video sequences and suffer from visual feature entanglement between robotic agents, manipulated objects, and environmental contexts, thereby limiting semantic decoupling capability for atomic actions and compromising model generalizability.To overcome these challenges, this work presents RoboAct-CLIP with dual technical contributions: 1) A dataset reconstruction framework that performs semantic-constrained action unit segmentation and re-annotation on open-source robotic videos, constructing purified training sets containing singular atomic actions (e.g., "grasp"); 2) A temporal-decoupling fine-tuning strategy based on Contrastive Language-Image Pretraining (CLIP) architecture, which disentangles temporal action features across video frames from object-centric characteristics to achieve hierarchical representation learning of robotic atomic actions.Experimental results in simulated environments demonstrate that the RoboAct-CLIP pretrained model achieves a 12% higher success rate than baseline VLMs, along with superior generalization in multi-object manipulation tasks.

📄 PDF Abstract BibTeX arXiv:2504.02069

Code (0)

등록된 구현이 없습니다.

Tasks

Action UnderstandingRepresentation Learning

Similar Papers 제목 키워드 기반

AVA: A Video Dataset of Spatio-temporally Localized Atomic Visual Actions

2017-05-23 · CVPR 2018 6 · Chunhui Gu, Chen Sun, David A. Ross, Carl Vondrick 외

This paper introduces a video dataset of spatio-temporally localized Atomic Visual Actions (AVA). The AVA dataset densely annotates 80 atomic visual actions in 430 15-minute video clips, where actions are localized in sp…

Actin DetectionAction DetectionAction LocalizationAction Recognition+3

Video ReCap: Recursive Captioning of Hour-Long Videos

2024-02-20 · CVPR 2024 1 · Md Mohaiminul Islam, Ngan Ho, Xitong Yang, Tushar Nagarajan 외

Most video captioning models are designed to process short video clips of few seconds and output text describing low-level visual concepts (e.g., objects, scenes, atomic actions). However, most real-world videos last for…

EgoSchemaVideo CaptioningVideo UnderstandingZero-Shot Video Question Answer

Temporally Consistent and Controllable Video Generation of 2D Cine CMR via Latent Space Motion Modeling

2026-06-08 · Yiheng Cao, Gustavo Andrade-Miranda, Jiatian Zhang, Guillaume Sallé 외 arxiv

Cine cardiac magnetic resonance is the gold standard for assessing cardiac function, but the scarcity of public datasets limits the development of advanced data-driven models. To address this limitation, we propose a gen…

Video Generation

Video Event Recognition and Anomaly Detection by Combining Gaussian Process and Hierarchical Dirichlet Process Models

2018-02-09 · Michael Ying Yang, Wentong Liao, Yanpeng Cao, Bodo Rosenhahn

In this paper, we present an unsupervised learning framework for analyzing activities and interactions in surveillance videos. In our framework, three levels of video events are connected by Hierarchical Dirichlet Proces…

Anomaly DetectionGeneral Classification

CLIP-Driven Universal Model for Organ Segmentation and Tumor Detection

2023-01-02 · ICCV 2023 1 · Jie Liu, Yixiao Zhang, Jie-Neng Chen, Junfei Xiao 외

An increasing number of public datasets have shown a marked impact on automated organ segmentation and tumor detection. However, due to the small size and partially labeled problem of each dataset, as well as a limited i…

Organ SegmentationSegmentationTransfer Learning