paper-with-me

홈 › Papers

Autoregressive Universal Video Segmentation Model

2025-08-26 · Miran Heo, Sukjun Hwang, Min-Hung Chen, Yu-Chiang Frank Wang, Albert Gu, Seon Joo Kim, Ryo Hachiuma arxiv

Recent video foundation models such as SAM2 excel at prompted video segmentation by treating masks as a general-purpose primitive. However, many real-world settings require unprompted segmentation that aims to detect and track all objects in a video without external cues, leaving today's landscape fragmented across task-specific models and pipelines. We recast streaming video segmentation as sequential mask prediction, analogous to language modeling, and introduce the Autoregressive Universal Segmentation Model (AUSM), a single architecture that unifies both prompted and unprompted video segmentation. Built on recent state-space models, AUSM maintains a fixed-size spatial state and scales to video streams of arbitrary length. Furthermore, all components of AUSM are designed for parallel training across frames, yielding substantial speedups over iterative training. On standard benchmarks (DAVIS17, YouTube-VOS 2018 & 2019, MOSE, YouTube-VIS 2019 & 2021, and OVIS) AUSM outperforms prior universal streaming video segmentation methods and achieves up to 2.5x faster training on 16-frame sequences.

📄 PDF Abstract BibTeX arXiv:2508.19242

Code (0)

등록된 구현이 없습니다.

Tasks

Video Segmentation

Similar Papers 제목 키워드 기반

Mask2Former for Video Instance Segmentation

2021-12-20 · Bowen Cheng, Anwesa Choudhuri, Ishan Misra, Alexander Kirillov 외

We find Mask2Former also achieves state-of-the-art performance on video instance segmentation without modifying the architecture, the loss or even the training pipeline. In this report, we show universal image segmentati…

Image SegmentationInstance SegmentationPanoptic SegmentationSegmentation+4

HyperSeg: Towards Universal Visual Segmentation with Large Language Model

2024-11-26 · Cong Wei, Yujie Zhong, Haoxian Tan, Yong liu 외

This paper aims to address universal segmentation for image and video perception with the strong reasoning ability empowered by Visual Large Language Models (VLLMs). Despite significant progress in current unified segmen…

Language ModelingLarge Language ModelOpen Vocabulary Semantic SegmentationPanoptic Segmentation+9

DVIS++: Improved Decoupled Framework for Universal Video Segmentation

2023-12-20 · Tao Zhang, Xingye Tian, Yikang Zhou, Shunping Ji 외

We present the \textbf{D}ecoupled \textbf{VI}deo \textbf{S}egmentation (DVIS) framework, a novel approach for the challenging task of universal video segmentation, including video instance segmentation (VIS), video seman…

Contrastive LearningDenoisingInstance SegmentationPanoptic Segmentation+6

Autoregressive Video Generation beyond Next Frames Prediction

2025-09-28 · Sucheng Ren, Chen Chen, Zhenbang Wang, Liangchen Song 외 arxiv

Autoregressive models for video generation typically operate frame-by-frame, extending next-token prediction from language to video's temporal dimension. We question that unlike word as token is universally agreed in lan…

Video Generation

HyperSeg: Hybrid Segmentation Assistant with Fine-grained Visual Perceiver

2025-01-01 · CVPR 2025 1 · Cong Wei, Yujie Zhong, Haoxian Tan, Yong liu 외

This paper aims to address universal segmentation for image and video perception with the strong reasoning ability empowered by Visual Large Language Models (VLLMs). Despite significant progress in current unified se…

Reasoning SegmentationSegmentationUniversal SegmentationVideo Segmentation+2