MotionBits: Video Segmentation through Motion-Level Analysis of Rigid Bodies
Rigid bodies constitute the smallest manipulable elements in the real world, and understanding how they physically interact is fundamental to embodied reasoning and robotic manipulation. Thus, accurate detection, segmentation, and tracking of moving rigid bodies is essential for enabling reasoning modules to interpret and act in diverse environments. However, current segmentation models trained on semantic grouping are limited in their ability to provide meaningful interaction-level cues for completing embodied tasks. To address this gap, we introduce MotionBit, a novel concept that, unlike prior formulations, defines the smallest unit in motion-based segmentation through kinematic spatial twist equivalence, independent of semantics. In this paper, we contribute (1) the MotionBit concept and definition, (2) a hand-labeled benchmark, called MoRiBo, for evaluating moving rigid-body segmentation across robotic manipulation and human-in-the-wild videos, and (3) a learning-free graph-based MotionBits segmentation method that outperforms state-of-the-art embodied perception methods by 37.3\% in macro-averaged mIoU on the MoRiBo benchmark. Finally, we demonstrate the effectiveness of MotionBits segmentation for downstream embodied reasoning and manipulation tasks, highlighting its importance as a fundamental primitive for understanding physical interactions.
Code (0)
등록된 구현이 없습니다.
Tasks
Video SegmentationSimilar Papers 제목 키워드 기반
InstMove: Instance Motion for Object-centric Video Segmentation
Despite significant efforts, cutting-edge video segmentation methods still remain sensitive to occlusion and rapid movement, due to their reliance on the appearance of objects in the form of object embeddings, which are …
ObjectOptical Flow EstimationSegmentationVideo Segmentation+1FusionSeg: Learning to Combine Motion and Appearance for Fully Automatic Segmentation of Generic Objects in Videos
We propose an end-to-end learning framework for segmenting generic objects in videos. Our method learns to combine appearance and motion information to produce pixel level segmentation masks for all prominent objects in …
SegmentationStructured PredictionVideo SegmentationVideo Semantic SegmentationObject Detection, Tracking, and Motion Segmentation for Object-level Video Segmentation
We present an approach for object segmentation in videos that combines frame-level object detection with concepts from object tracking and motion segmentation. The approach extracts temporally consistent object tubes bas…
Motion SegmentationObjectobject-detectionObject Detection+5FusionSeg: Learning to combine motion and appearance for fully automatic segmention of generic objects in videos
We propose an end-to-end learning framework for segmenting generic objects in videos. Our method learns to combine appearance and motion information to produce pixel level segmentation masks for all prominent objects in …
SegmentationStructured PredictionUnsupervised Video Object SegmentationVideo Segmentation+1Decoupled Motion Expression Video Segmentation
Motion expression video segmentation aims to segment objects based on input motion descriptions. Compared with traditional referring video object segmentation, it focuses on motion and multi-object expressions and is…
Instance SegmentationReferring Video Object SegmentationSegmentationSemantic Segmentation+4