paper-with-me

홈 › Papers

MOBIUS: Big-to-Mobile Universal Instance Segmentation via Multi-modal Bottleneck Fusion and Calibrated Decoder Pruning

2025-10-16 · Mattia Segu, Marta Tintore Gazulla, Yongqin Xian, Luc Van Gool, Federico Tombari arxiv

Scaling up model size and training data has advanced foundation models for instance-level perception, achieving state-of-the-art in-domain and zero-shot performance across object detection and segmentation. However, their high computational cost limits adoption on resource-constrained platforms. We first examine the limitations of existing architectures in enabling efficient edge deployment without compromising performance. We then introduce MOBIUS, a family of foundation models for universal instance segmentation, designed for Pareto-optimal downscaling to support deployment across devices ranging from high-end accelerators to mobile hardware. To reduce training and inference demands, we propose: (i) a bottleneck pixel decoder for efficient multi-scale and multi-modal fusion, (ii) a language-guided uncertainty calibration loss for adaptive decoder pruning, and (iii) a streamlined, unified training strategy. Unlike efficient baselines that trade accuracy for reduced complexity, MOBIUS reduces pixel and transformer decoder FLOPs by up to 55% and 75%, respectively, while maintaining state-of-the-art performance in just a third of the training iterations. MOBIUS establishes a new benchmark for efficient segmentation on both high-performance computing platforms and mobile devices.

📄 PDF Abstract BibTeX arXiv:2510.15026

Code (0)

등록된 구현이 없습니다.

Tasks

Instance SegmentationObject Detection

Similar Papers 제목 키워드 기반

Segmenting, Fast and Slow: Real-Time Open-Vocabulary Video Instance Segmentation with Dual-Path Processing

2026-06-30 · Luca Barsellotti, Martin Sundermeyer, Mattia Segu, Nikita Araslanov 외 arxiv

Object-centric models inspired by DETR have become the dominant paradigm for open-vocabulary video instance segmentation (OV-VIS). While recent efforts have reduced the computational cost of pixel decoding, textual modal…

Video Instance Segmentation

MOBIUS: A Multi-Modal Bipedal Robot that can Walk, Crawl, Climb, and Roll

2025-11-03 · Alexander Schperberg, Yusuke Tanaka, Stefano Di Cairano, Dennis Hong arxiv

This paper presents the MOBIUS platform, a bipedal robot capable of walking, crawling, climbing, and rolling. MOBIUS features four limbs, two 6-DoF arms with two-finger grippers for manipulation and climbing, and two 4-D…

Reinforcement Learning

Shear-Free Viewport Magnification for 360-Degree via Spherical Mobius Boosts

2026-06-14 · Boyang Li, Hezhao Xu arxiv

Viewport-adaptive 360-degree imaging seeks to allocate a fixed sampling budget to the region a viewer is likely to observe. Existing view-biased projections increase viewport resolution through non-conformal warps, which…

MobileInst: Video Instance Segmentation on the Mobile

2023-03-30 · Renhong Zhang, Tianheng Cheng, Shusheng Yang, Haoyi Jiang 외

Video instance segmentation on mobile devices is an important yet very challenging edge AI problem. It mainly suffers from (1) heavy computation and memory costs for frame-by-frame pixel-level instance perception and (2)…

CPUDecoderInstance SegmentationSegmentation+2

Computational Aspects of the Mobius Transform

2013-03-27 · Robert Kennes, Philippe Smets

In this paper we associate with every (directed) graph G a transformation called the Mobius transformation of the graph G. The Mobius transformation of the graph (O) is of major significance for Dempster-Shafer theory of…