paper-with-me

홈 › Papers

Omni-Scan: Creating Visually-Accurate Digital Twin Object Models Using a Bimanual Robot with Handover and Gaussian Splat Merging

2025-08-01 · Tianshuang Qiu, Zehan Ma, Karim El-Refai, Hiya Shah, Chung Min Kim, Justin Kerr, Ken Goldberg arxiv

3D Gaussian Splats (3DGSs) are 3D object models derived from multi-view images. Such "digital twins" are useful for simulations, virtual reality, marketing, robot policy fine-tuning, and part inspection. 3D object scanning usually requires multi-camera arrays, precise laser scanners, or robot wrist-mounted cameras, which have restricted workspaces. We propose Omni-Scan, a pipeline for producing high-quality 3D Gaussian Splat models using a bi-manual robot that grasps an object with one gripper and rotates the object with respect to a stationary camera. The object is then re-grasped by a second gripper to expose surfaces that were occluded by the first gripper. We present the Omni-Scan robot pipeline using DepthAny-thing, Segment Anything, as well as RAFT optical flow models to identify and isolate objects held by a robot gripper while removing the gripper and the background. We then modify the 3DGS training pipeline to support concatenated datasets with gripper occlusion, producing an omni-directional (360 degree view) model of the object. We apply Omni-Scan to part defect inspection, finding that it can identify visual or geometric defects in 12 different industrial and household objects with an average accuracy of 83%. Interactive videos of Omni-Scan 3DGS models can be found at https://berkeleyautomation.github.io/omni-scan/

📄 PDF Abstract BibTeX arXiv:2508.00354

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OmniRe: Omni Urban Scene Reconstruction

2024-08-29 · Ziyu Chen, Jiawei Yang, Jiahui Huang, Riccardo de Lutio 외

We introduce OmniRe, a comprehensive system for efficiently creating high-fidelity digital twins of dynamic real-world scenes from on-device logs. Recent methods using neural fields or Gaussian Splatting primarily focus …

3DGS

Real5-OmniDocBench: A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild

2026-03-04 · Changda Zhou, Ziyue Gao, Xueqing Wang, Tingquan Gao 외 arxiv

While Vision-Language Models (VLMs) achieve near-perfect scores on digital document benchmarks like OmniDocBench, their performance in the unpredictable physical world remains largely unknown due to the lack of controlle…

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation

2026-05-12 · Che Liu, Lichao Ma, Xiangyu Tony Zhang, Yuxin Zhang 외 arxiv

Omni-modal language models are intended to jointly understand audio, visual inputs, and language, but benchmark gains can be inflated when visual evidence alone is enough to answer a query. We study whether current omni-…

OmniMamba4D: Spatio-temporal Mamba for longitudinal CT lesion segmentation

2025-04-13 · Justin Namuk Kim, Yiqiao Liu, Rajath Soans, Keith Persson 외

Accurate segmentation of longitudinal CT scans is important for monitoring tumor progression and evaluating treatment responses. However, existing 3D segmentation models solely focus on spatial information. To address th…

Computational EfficiencyLesion SegmentationMambaSegmentation

Whole-Body Image-to-Image Translation for a Virtual Scanner in a Healthcare Digital Twin

2025-03-18 · Valerio Guarrasi, Francesco Di Feola, Rebecca Restivo, Lorenzo Tronchin 외

Generating positron emission tomography (PET) images from computed tomography (CT) scans via deep learning offers a promising pathway to reduce radiation exposure and costs associated with PET imaging, improving patient …

Computed Tomography (CT)Image GenerationImage-to-Image TranslationTranslation