paper-with-me

Papers

MEPA: Multi-Scale Representation Alignment for Visual Autoregressive Modeling with Mixture of Experts

2026-07-01 · Nuoyan Zhou, Zhijun Tu, Lei Yu, Kun Cheng, Jie Hu, Nannan Wang, Xinghao Chen arxiv

Visual AutoRegressive modeling (VAR) has pioneered a coarse-to-fine multi-scale autoregressive generative paradigm, demonstrating strong capabilities in image generation. However, VAR still suffers from inherent deficiencies in multi-scale representation learning. Specifically, lower scales primarily capture global semantics, while higher scales focus on fine-grained details. Employing a shared architecture across scales induces optimization conflicts. Moreover, due to the causal autoregressive process, inaccurate semantics at early scales can propagate and significantly degrade the final output. To address these issues, we introduce a scale-aware token-routed Mixture of Experts (MoE) architecture, allowing scale-adaptive expert selection, thereby facilitating decoupled representation learning across scales. In addition, we enhance semantic modeling at early scales by incorporating external self-supervised features. Unlike naive alignment, we analyse and design a residual feature aggregation scheme tailored to the VAR paradigm. Extensive experiments show that our method significantly improves both training efficiency and generation quality. On the ImageNet 256*256 benchmark, our model achieves a superior FID compared to the dense baseline while requiring only half of the default training epochs and a smaller parameter budget, with a merely marginal increase in training cost. Moreover, the performance gap further widens with larger training epochs.

📄 PDF Abstract BibTeX arXiv:2607.00371

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningImage Generation

Similar Papers 제목 키워드 기반

Homepage2Vec: Language-Agnostic Website Embedding and Classification

2022-01-10 · Sylvain Lugeon, Tiziano Piccardi, Robert West

Currently, publicly available models for website classification do not offer an embedding method and have limited support for languages beyond English. We release a dataset of more than two million category-labeled websi…

Classification

USD: A User-Intent-Driven Sampling and Dual-Debiasing Framework for Large-Scale Homepage Recommendations

2025-07-09 · Jiaqi Zheng, Cheng Guo, Yi Cao, Chaoqun Hou 외

Large-scale homepage recommendations face critical challenges from pseudo-negative samples caused by exposure bias, where non-clicks may indicate inattention rather than disinterest. Existing work lacks thorough analysis…

Marketing

Any3D-VLA: Enhancing VLA Robustness via Diverse Point Clouds

2026-01-31 · Xianzhe Fan, Shengliang Deng, Xiaoyang Wu, Yuxiang Lu 외 arxiv

Existing Vision-Language-Action (VLA) models typically take 2D images as visual input, which limits their spatial understanding in complex scenes. How can we incorporate 3D information to enhance VLA capabilities? We con…

Point Clouds

Layered Image Vectorization via Semantic Simplification

2024-06-08 · CVPR 2025 1 · Zhenyu Wang, Jianxi Huang, Zhida Sun, Yuanhao Gong 외

This work presents a progressive image vectorization technique that reconstructs the raster image as layer-wise vectors from semantic-aligned macro structures to finer details. Our approach introduces a new image simplif…

Semantic Segmentation

A Hierarchical Representation Network for Accurate and Detailed Face Reconstruction from In-The-Wild Images

2023-02-28 · CVPR 2023 1 · Biwen Lei, Jianqiang Ren, Mengyang Feng, Miaomiao Cui 외

Limited by the nature of the low-dimensional representational capacity of 3DMM, most of the 3DMM-based face reconstruction (FR) methods fail to recover high-frequency facial details, such as wrinkles, dimples, etc. Some …

3D Face ReconstructionDisentanglementFace Reconstruction