paper-with-me

Papers

Mask2Real-WM: Segmentation Masks as a Sim-to-Real Bridge for Controllable Dexterous World Models

2026-07-05 · Riccardo O. Feingold, Davide Liconti, Chenyu Yang, Robert K. Katzschmann arxiv

Action-conditioned world models allow robots to predict the future consequences of candidate actions without additional physical interaction, supporting policy evaluation, planning, and data augmentation. We present Mask2Real-WM, a two-stage action-conditioned world model for dexterous manipulation that decouples pixel prediction into a dynamics model and a rendering model. The dynamics model predicts future segmentation masks from past masks and 23-DoF action sequences. The rendering model maps the predicted masks to photorealistic RGB using a ControlNet-augmented Stable Video Diffusion backbone. The smaller sim-to-real gap in segmentation space enables the dynamics model to benefit from large-scale pretraining on over 50 h of synthetic simulation data, followed by fine-tuning on fewer than 2.5 h of real demonstrations. Experiments on a dexterous pick-and-place benchmark show that mask conditioning and simulation pretraining are both required for per-DoF action controllability across all 23 degrees of freedom. In contrast, monolithic baselines capture broad hand and end-effector trajectories but do not reliably reflect fine-grained, per-joint action effects.

📄 PDF Abstract BibTeX arXiv:2607.04546

Code (0)

등록된 구현이 없습니다.

Tasks

Data Augmentation

Similar Papers 제목 키워드 기반

SEMI-PointRend: Improved Semiconductor Wafer Defect Classification and Segmentation as Rendering

2023-02-19 · MinJin Hwang, Bappaditya Dey, Enrique Dehaerne, Sandip Halder 외

In this study, we applied the PointRend (Point-based Rendering) method to semiconductor defect segmentation. PointRend is an iterative segmentation algorithm inspired by image rendering in computer graphics, a new image …

Image SegmentationInstance SegmentationSegmentationSemantic Segmentation

BoxTeacher: Exploring High-Quality Pseudo Labels for Weakly Supervised Instance Segmentation

2022-10-11 · CVPR 2023 1 · Tianheng Cheng, Xinggang Wang, Shaoyu Chen, Qian Zhang 외

Labeling objects with pixel-wise segmentation requires a huge amount of human labor compared to bounding boxes. Most existing methods for weakly supervised instance segmentation focus on designing heuristic losses with p…

Box-supervised Instance SegmentationInstance SegmentationSegmentationSemantic Segmentation+2

Bridging the Domain Gap: Self-Supervised 3D Scene Understanding with Foundation Models

2023-05-15 · NeurIPS 2023 11 · Zhimin Chen, Longlong Jing, Yingwei Li, Bing Li

Foundation models have achieved remarkable results in 2D and language tasks like image segmentation, object detection, and visual-language understanding. However, their potential to enrich 3D scene representation learnin…

3D Object DetectionImage CaptioningImage SegmentationKnowledge Distillation+6

GenMask: Adapting DiT for Segmentation via Direct Mask Generation

2026-03-25 · Yuhuan Yang, Xianwei Zhuang, Yuxuan Cai, Chaofan Ma 외 arxiv

Recent approaches for segmentation have leveraged pretrained generative models as feature extractors, treating segmentation as a downstream adaptation task via indirect feature retrieval. This implicit use suffers from a…

Image Generation

Realistic Hair Simulation Using Image Blending

2019-04-19 · Mohamed Attia, Mohammed Hossny, Saeid Nahavandi, Anousha Yazdabadi 외

In this presented work, we propose a realistic hair simulator using image blending for dermoscopic images. This hair simulator can be used for benchmarking and validation of the hair removal methods and in data augmentat…

BenchmarkingData AugmentationDiagnostic