paper-with-me

홈 › Papers

StraIT: Non-autoregressive Generation with Stratified Image Transformer

2023-03-01 · Shengju Qian, Huiwen Chang, Yuanzhen Li, Zizhao Zhang, Jiaya Jia, Han Zhang

We propose Stratified Image Transformer(StraIT), a pure non-autoregressive(NAR) generative model that demonstrates superiority in high-quality image synthesis over existing autoregressive(AR) and diffusion models(DMs). In contrast to the under-exploitation of visual characteristics in existing vision tokenizer, we leverage the hierarchical nature of images to encode visual tokens into stratified levels with emergent properties. Through the proposed image stratification that obtains an interlinked token pair, we alleviate the modeling difficulty and lift the generative power of NAR models. Our experiments demonstrate that StraIT significantly improves NAR generation and out-performs existing DMs and AR methods while being order-of-magnitude faster, achieving FID scores of 3.96 at 256*256 resolution on ImageNet without leveraging any guidance in sampling or auxiliary image classifiers. When equipped with classifier-free guidance, our method achieves an FID of 3.36 and IS of 259.3. In addition, we illustrate the decoupled modeling process of StraIT generation, showing its compelling properties on applications including domain transfer.

📄 PDF Abstract BibTeX arXiv:2303.00750

Code (0)

등록된 구현이 없습니다.

Tasks

Image Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Machine Learning-Based Classification of Vessel Types in Straits Using AIS Tracks

2025-09-09 · Jonatan Katz Nielsen arxiv

Accurate recognition of vessel types from Automatic Identification System (AIS) tracks is essential for safety oversight and combating illegal, unreported, and unregulated (IUU) activity. This paper presents a strait-sca…

Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation

2025-09-23 · Shufan Li, Jiuxiang Gu, Kangning Liu, Zhe Lin 외 arxiv

We propose Lavida-O, a unified Masked Diffusion Model (MDM) for multimodal understanding and generation. Unlike existing multimodal MDMs such as MMaDa and Muddit which only support simple image-level understanding tasks …

Text-to-Image GenerationMultimodal ReasoningImage Editing

Text Generation: Reexamining G-TAG with Abstract Categorial Grammars (G\'en\'eration de textes : G-TAG revisit\'e avec les Grammaires Cat\'egorielles Abstraites) [in French]

2014-07-01 · JEPTALNRECITAL 2014 7 · Laurence Danlos, Aleks Maskharashvili, re, Sylvain Pogodalla
TAGText Generation

Improved Masked Image Generation with Token-Critic

2022-09-09 · José Lezama, Huiwen Chang, Lu Jiang, Irfan Essa

Non-autoregressive generative transformers recently demonstrated impressive image generation performance, and orders of magnitude faster sampling than their autoregressive counterparts. However, optimal parallel sampling…

DiversityImage Generation

Head-Aware Key-Value Compression for Efficient Autoregressive Image Generation

2026-05-20 · Guotao Liang, Baoquan Zhang, Zhiyuan Wen, Yunming Ye arxiv

Autoregressive (AR) visual generation has achieved remarkable performance but suffers from high memory usage and low throughput, as it requires caching previously generated visual tokens. Recent research has shown that r…

Image Generation