paper-with-me

홈 › Papers

Cross-scale Aligned Supervision for Training GANs

2026-05-26 · Sangeek Hyun, MinKyu Lee, Jae-Pil Heo arxiv

Modern GANs often introduce adversarial supervision on intermediate generator outputs and interpret the resulting multi-stage synthesis as coarse-to-fine hierarchical generation. In this work, we challenge this interpretation. We argue that standard scale-wise adversarial supervision does not construct a proper coarse-to-fine hierarchy: each intermediate image is independently pushed toward the real distribution at its own resolution, but this scale-wise realism does not ensure that outputs across stages represent the identical generated sample. Moreover, the scale-specific image produced at each stage is not used as an explicit refinement target for the subsequent stage. Therefore, its adversarial loss can improve a scale-specific output without constraining later stages to preserve the same sample trajectory, allowing them to move toward a different sample rather than refine the previous output. We refer to this problem as a cross-scale trajectory misalignment problem. To resolve it, we propose CAT, a Cross-scale Aligned Transformer for multi-scale adversarial generation. CAT keeps the discriminator scale-wise, so each intermediate output is evaluated at its own resolution, while adding a simple generator-side consistency regularization that aligns intermediate outputs with the final output. On class-conditional ImageNet-256, CAT-H/2 achieves an FID-50K of 1.56 with one-step inference after only 60 training epochs, outperforming strong one-step GAN and diffusion/flow baselines.

📄 PDF Abstract BibTeX arXiv:2605.26449

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Classification-Based Perspective on GAN Distributions

2018-01-01 · ICLR 2018 1 · Shibani Santurkar, Ludwig Schmidt, Aleksander Madry

A fundamental, and still largely unanswered, question in the context of Generative Adversarial Networks (GANs) is whether GANs are actually able to capture the key characteristics of the datasets they are trained on. The…

ClassificationDiversityGeneral Classification

Scalable GANs with Transformers

2025-09-29 · Sangeek Hyun, MinKyu Lee, Jae-Pil Heo arxiv

Scalability has driven recent advances in generative modeling, yet its principles remain underexplored for adversarial learning. We investigate the scalability of Generative Adversarial Networks (GANs) through two design…

Self-Supervised GANs via Auxiliary Rotation Loss

2018-11-27 · CVPR 2019 6 · Ting Chen, Xiaohua Zhai, Marvin Ritter, Mario Lucic 외

Conditional GANs are at the forefront of natural image synthesis. The main drawback of such models is the necessity for labeled data. In this work we exploit two popular unsupervised learning techniques, adversarial trai…

Image GenerationRepresentation Learning

PISCES: Annotation-free Text-to-Video Post-Training via Optimal Transport-Aligned Rewards

2026-02-02 · Minh-Quan Le, Gaurav Mittal, Cheng Zhao, David Gu 외 arxiv

Text-to-video (T2V) generation aims to synthesize videos with high visual quality and temporal consistency that are semantically aligned with input text. Reward-based post-training has emerged as a promising direction to…

Reinforcement LearningVideo Generation

GreenRFM: Learning a resource-efficient radiology vision-language foundation model via supervision-centric pre-training

2026-03-06 · Yingtai Li, Shuai Ming, Qiuli Wang, Mingyue Zhao 외 arxiv

Radiology foundation models (RFMs) have largely inherited the scale-first recipe of natural-image vision--language pre-training. This recipe is difficult to deploy in 3D radiology, where training corpora are smaller, rep…