paper-with-me

Papers

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder

2026-06-09 · Yitong Chen, Zijie Diao, Junke Wang, Lingyu Kong, Yixuan Ren, Bo He, Yu-Gang Jiang, Zuxuan Wu arxiv

Built on pretrained vision foundation models (VFMs), representation autoencoders (RAEs) have recently emerged as a promising approach for constructing semantically rich latent spaces for image generation. However, their reconstruction quality often remains suboptimal, largely because deep VFM representations do not preserve sufficient fine-grained visual detail. This limitation becomes even more severe after discretization, where missing low-level information is difficult to recover. In fact, we observe that shallow VFM features retain considerably richer local appearance and structural detail, which complements the high-level semantics carried by deep features used in existing RAEs. Motivated by this complementary property, we propose Ideal, an In-depth Alignment framework for discrete representation autoencoding. By jointly aligning quantized tokens with both shallow and deep VFM features, Ideal enables the resulting discrete visual tokens to preserve both visual fidelity and rich semantics. Extensive experiments demonstrate that Ideal yields superior reconstruction performance, achieving 0.61 rFID on ImageNet and outperforming the previous best method by 0.28. When used for autoregressive image generation, Ideal further produces a gFID of 1.89, establishing a new state of the art for autoregressive image generation.

📄 PDF Abstract BibTeX arXiv:2606.11096

Code (0)

등록된 구현이 없습니다.

Tasks

Image Generation

Similar Papers 제목 키워드 기반

Modelling and analysis of the 8 filters from the "master key filters hypothesis" for depthwise-separable deep networks in relation to idealized receptive fields based on scale-space theory

2025-09-16 · Tony Lindeberg, Zahra Babaiee, Peyman M. Kiasari arxiv

This paper presents the results of analysing and modelling a set of 8 ``master key filters'', which have been extracted by applying a clustering approach to the receptive fields learned in depthwise-separable deep networ…

Learnable Chernoff Baselines for Inference-Time Alignment

2026-02-08 · Sunil Madhow, Yuchen Liang, Ness Shroff, Yingbin Liang 외 arxiv

We study inference-time reward-guided alignment for generative models. Existing methods often rely on either architecture-specific adaptations or computationally costly inference procedures. We introduce Learnable Cherno…

C2PD: Continuity-Constrained Pixelwise Deformation for Guided Depth Super-Resolution

2025-01-13 · Jiahui Kang, Qing Cai, Runqing Tan, Yimei Liu 외

Guided depth super-resolution (GDSR) has demonstrated impressive performance across a wide range of domains, with numerous methods being proposed. However, existing methods often treat depth maps as images, where shading…

Super-Resolution

Learning Pixel-wise Continuous Depth Representation via Clustering for Depth Completion

2024-02-21 · Chen Shenglun, Zhang Hong, Ma XinZhu, Wang Zhihui 외

Depth completion is a long-standing challenge in computer vision, where classification-based methods have made tremendous progress in recent years. However, most existing classification-based methods rely on pre-defined …

ClusteringDepth Completion

SD-6DoF-ICLK: Sparse and Deep Inverse Compositional Lucas-Kanade Algorithm on SE(3)

2021-03-30 · Timo Hinzmann, Roland Siegwart

This paper introduces SD-6DoF-ICLK, a learning-based Inverse Compositional Lucas-Kanade (ICLK) pipeline that uses sparse depth information to optimize the relative pose that best aligns two images on SE(3). To compute th…

Simultaneous Localization and Mapping