paper-with-me

홈 › Papers

RPiAE: A Representation-Pivoted Autoencoder Enhancing Both Image Generation and Editing

2026-03-19 · Yue Gong, Hongyu Li, Shanyuan Liu, Bo Cheng, Yuhang Ma, Liebucha Wu, Xiaoyu Wu, Manyuan Zhang, Dawei Leng, Yuhui Yin, Lijun Zhang arxiv

Diffusion models have become the dominant paradigm for image generation and editing, with latent diffusion models shifting denoising to a compact latent space for efficiency and scalability. Recent attempts to leverage pretrained visual representation models as tokenizer priors either align diffusion features to representation features or directly reuse representation encoders as frozen tokenizers. Although such approaches can improve generation metrics, they often suffer from limited reconstruction fidelity due to frozen encoders, which in turn degrades editing quality, as well as overly high-dimensional latents that make diffusion modeling difficult. To address these limitations, We propose Representation-Pivoted AutoEncoder, a representation-based tokenizer that improves both generation and editing. We introduce Representation-Pivot Regularization, a training strategy that enables a representation-initialized encoder to be fine-tuned for reconstruction while preserving the semantic structure of the pretrained representation space, followed by a variational bridge which compress latent space into a compact one for better diffusion modeling. We adopt an objective-decoupled stage-wise training strategy that sequentially optimizes generative tractability and reconstruction-fidelity objectives. Together, these components yield a tokenizer that preserves strong semantics, reconstructs faithfully, and produces latents with reduced diffusion modeling complexity. Experiments demonstrate that RPiAE outperforms other visual tokenizers on text-to-image generation and image editing, while delivering the best reconstruction fidelity among representation-based tokenizers.

📄 PDF Abstract BibTeX arXiv:2603.19206

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Image GenerationImage Editing

Similar Papers 제목 키워드 기반

Kernel Quadrature with Randomly Pivoted Cholesky

2023-09-21 · NeurIPS 2023 11

This paper presents new quadrature rules for functions in a reproducing kernel Hilbert space using nodes drawn by a sampling algorithm known as randomly pivoted Cholesky. The resulting computational procedure compares fa…

Kernel quadrature with randomly pivoted Cholesky

2023-06-06 · Ethan N. Epperly, Elvira Moreno

This paper presents new quadrature rules for functions in a reproducing kernel Hilbert space using nodes drawn by a sampling algorithm known as randomly pivoted Cholesky. The resulting computational procedure compares fa…

EVOQUER: Enhancing Temporal Grounding with Video-Pivoted BackQuery Generation

2021-09-10 · Yanjun Gao, Lulu Liu, Jason Wang, Xin Chen 외

Temporal grounding aims to predict a time interval of a video clip corresponding to a natural language query input. In this work, we present EVOQUER, a temporal grounding framework incorporating an existing text-to-video…

TranslationVideo Grounding

Unifying Heterogeneous Multi-Modal Remote Sensing Detection Via Language-Pivoted Pretraining

2026-03-02 · Yuxuan Li, Yuming Chen, Yunheng Li, Ming-Ming Cheng 외 arxiv

Heterogeneous multi-modal remote sensing object detection aims to accurately detect objects from diverse sensors (e.g., RGB, SAR, Infrared). Existing approaches largely adopt a late alignment paradigm, in which modality …

Object Detection

From Words to Sentences: A Progressive Learning Approach for Zero-resource Machine Translation with Visual Pivots

2019-06-03 · Shizhe Chen, Qin Jin, Jianlong Fu

The neural machine translation model has suffered from the lack of large-scale parallel corpora. In contrast, we humans can learn multi-lingual translations even without parallel texts by referring our languages to the e…

Machine TranslationSentenceTranslationWord Translation