paper-with-me

Papers

DPAR: Dynamic Patchification for Efficient Autoregressive Visual Generation

2025-12-26 · Divyansh Srivastava, Akshay Mehra, Pranav Maneriker, Debopam Sanyal, Vishnu Raj, Vijay Kamarshi, Fan Du, Joshua Kimball arxiv

Decoder-only autoregressive image generation typically relies on fixed-length tokenization schemes whose token counts grow quadratically with resolution, substantially increasing the computational and memory demands of attention. We present DPAR, a novel decoder-only autoregressive model that dynamically aggregates image tokens into a variable number of patches for efficient image generation. Our work is the first to demonstrate that next-token prediction entropy from a lightweight and unsupervised autoregressive model provides a reliable criterion for merging tokens into larger patches based on information content. DPAR makes minimal modifications to the standard decoder architecture, ensuring compatibility with multimodal generation frameworks and allocating more compute to generation of high-information image regions. Further, we demonstrate that training with dynamically sized patches yields representations that are robust to patch boundaries, allowing DPAR to scale to larger patch sizes at inference. DPAR reduces token count by 1.81x and 2.06x on Imagenet 256 and 384 generation resolution respectively, leading to a reduction of up to 40% FLOPs in training costs. Further, our method exhibits faster convergence and improves FID by up to 27.1% relative to baseline models.

📄 PDF Abstract BibTeX arXiv:2512.21867

Code (0)

등록된 구현이 없습니다.

Tasks

multimodal generationImage Generation

Similar Papers 제목 키워드 기반

dpart: Differentially Private Autoregressive Tabular, a General Framework for Synthetic Data Generation

2022-07-12 · Sofiane Mahiou, Kai Xu, Georgi Ganev

We propose a general, flexible, and scalable framework dpart, an open source Python library for differentially private synthetic data generation. Central to the approach is autoregressive modelling -- breaking the joint …

Synthetic Data Generation

Open-World Dynamic Prompt and Continual Visual Representation Learning

2024-09-09 · Youngeun Kim, Jun Fang, Qin Zhang, Zhaowei Cai 외

The open world is inherently dynamic, characterized by ever-evolving concepts and distributions. Continual learning (CL) in this dynamic open-world environment presents a significant challenge in effectively generalizing…

Continual LearningImage RetrievalPrompt LearningRepresentation Learning

Ms. Forcing: Efficient Streaming Video Generation with Multi-Scale Patchification and Attention

2026-07-23 · Zekun Li, Xiaoyan Cong, Hongyu Li, Zhiyang Dou 외 arxiv

Streaming video diffusion models have made substantial progress toward interactive and dynamic world simulation, but the nested autoregressive and denoising loops of conventional next-frame generation hinder real-time de…

Video Generation

dParallel: Learnable Parallel Decoding for dLLMs

2025-09-30 · Zigeng Chen, Gongfan Fang, Xinyin Ma, Ruonan Yu 외 arxiv

Diffusion large language models (dLLMs) have recently drawn considerable attention within the research community as a promising alternative to autoregressive generation, offering parallel token prediction and lower infer…

View about consumption tax and grandchildren

2021-02-09 · Eiji Yamamura

In Japan, the increase in the consumption tax rate, a measure of balanced public finance, reduces the inequality of fiscal burden between the present and future generations. This study estimates the effect of grandchildr…