paper-with-me

홈 › Papers

Progressive Supernet Training for Efficient Visual Autoregressive Modeling

2025-11-20 · Xiaoyue Chen, Yuling Shi, Kaiyuan Li, Huandong Wang, Yong Li, Xiaodong Gu, Xinlei Chen, Mingbao Lin arxiv

Visual Auto-Regressive (VAR) models significantly reduce inference steps through the "next-scale" prediction paradigm. However, progressive multi-scale generation incurs substantial memory overhead due to cumulative KV caching, limiting practical deployment. We observe a scale-depth asymmetric dependency in VAR: early scales exhibit extreme sensitivity to network depth, while later scales remain robust to depth reduction. Inspired by this, we propose VARiant: by equidistant sampling, we select multiple subnets ranging from 16 to 2 layers from the original 30-layer VAR-d30 network. Early scales are processed by the full network, while later scales utilize subnet. Subnet and the full network share weights, enabling flexible depth adjustment within a single model. However, weight sharing between subnet and the entire network can lead to optimization conflicts. To address this, we propose a progressive training strategy that breaks through the Pareto frontier of generation quality for both subnets and the full network under fixed-ratio training, achieving joint optimality. Experiments on ImageNet demonstrate that, compared to the pretrained VAR-d30 (FID 1.95), VARiant-d16 and VARiant-d8 achieve nearly equivalent quality (FID 2.05/2.12) while reducing memory consumption by 40-65%. VARiant-d2 achieves 3.5 times speedup and 80% memory reduction at moderate quality cost (FID 2.97). In terms of deployment, VARiant's single-model architecture supports zero-cost runtime depth switching and provides flexible deployment options from high quality to extreme efficiency, catering to diverse application scenarios.

📄 PDF Abstract BibTeX arXiv:2511.16546

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Improving Ranking Correlation of Supernet with Candidates Enhancement and Progressive Training

2021-08-12 · Ziwei Yang, Ruyi Zhang, Zhi Yang, Xubo Yang 외

One-shot neural architecture search (NAS) applies weight-sharing supernet to reduce the unaffordable computation overhead of automated architecture designing. However, the weight-sharing technique worsens the ranking con…

Neural Architecture Search

CAR: Controllable Autoregressive Modeling for Visual Generation

2024-10-07 · Ziyu Yao, Jialin Li, Yifeng Zhou, Yong liu 외

Controllable generation, which enables fine-grained control over generated outputs, has emerged as a critical focus in visual generative models. Currently, there are two primary technical approaches in visual generation:…

Neural Architecture Search using Progressive Evolution

2022-03-03 · Nilotpal Sinha, Kuan-Wen Chen

Vanilla neural architecture search using evolutionary algorithms (EA) involves evaluating each architecture by training it from scratch, which is extremely time-consuming. This can be reduced by using a supernet to estim…

Evolutionary AlgorithmsNeural Architecture Search

Neighboring Autoregressive Modeling for Efficient Visual Generation

2025-03-12 · Yefei He, Yuanyu He, Shaoxuan He, Feng Chen 외

Visual autoregressive models typically adhere to a raster-order ``next-token prediction" paradigm, which overlooks the spatial and temporal locality inherent in visual content. Specifically, visual tokens exhibit signifi…

Image GenerationText to Image GenerationText-to-Image GenerationVideo Generation

Mixed-precision Supernet Training from Vision Foundation Models using Low Rank Adapter

2024-03-29 · Yuiko Sakuma, Masakazu Yoshimura, Junji Otsuka, Atsushi Irie 외

Compression of large and performant vision foundation models (VFMs) into arbitrary bit-wise operations (BitOPs) allows their deployment on various hardware. We propose to fine-tune a VFM to a mixed-precision quantized su…

Neural Architecture Search