paper-with-me

홈 › Papers

Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer

2025-11-24 · Haoyu Wu, Jingyi Xu, Qiaomu Miao, Dimitris Samaras, Hieu Le arxiv

Rotary positional embeddings (RoPE) are widely used in diffusion transformers (DiTs) to encode spatial relationships, yet their behavior with mixed-resolution tokens remains underexplored. A natural approach is to rescale token positions from different resolutions into a unified coordinate system before attention, but we show this fails. Our analysis shows that with RoPE, the attention similarity score is a highly structured and periodic function of token distance, so rescaling distances across resolutions moves token pairs to different regions of this periodic function, leading to incorrect attention scores. Motivated by this, we introduce Phase-Aligned Mixed-Resolution Attention (PMA), a training-free mechanism that stabilizes mixed-resolution attention. PMA modifies the RoPE position mapping to enforce a consistent positional scale for every query-key pair, ensuring that relative distances are evaluated under a single reference scale. To further improve local coherence near resolution transitions, we incorporate a lightweight boundary refinement module that softly exchanges features across adjacent scales. Experiments on image and video diffusion models validate our analysis and demonstrate consistent improvements in visual fidelity and computational efficiency.

📄 PDF Abstract BibTeX arXiv:2511.19778

Code (0)

등록된 구현이 없습니다.

Tasks

Computational Efficiency

Similar Papers 제목 키워드 기반

PGP-DiffSR: Phase-Guided Progressive Pruning for Efficient Diffusion-based Image Super-Resolution

2025-12-02 · Zhongbao Yang, Jiangxin Dong, Yazhou Yao, Jinhui Tang 외 arxiv

Although diffusion-based models have achieved impressive results in image super-resolution, they often rely on large-scale backbones such as Stable Diffusion XL (SDXL) and Diffusion Transformers (DiT), which lead to exce…

Image Super-Resolution

NeuralRemaster: Phase-Preserving Diffusion for Structure-Aligned Generation

2025-12-04 · Yu Zeng, Charles Ochoa, Mingyuan Zhou, Vishal M. Patel 외 arxiv

Standard diffusion corrupts data using Gaussian noise whose Fourier coefficients have random magnitudes and random phases. While effective for unconditional or text-to-image generation, corrupting phase components destro…

Image-to-Image TranslationText-to-Image GenerationContinuous ControlVideo Generation

FRAMER: Frequency-Aligned Self-Distillation with Adaptive Modulation Leveraging Diffusion Priors for Real-World Image Super-Resolution

2025-12-01 · Seungho Choi, Jeahun Sung, Jihyong Oh arxiv

Real-image super-resolution (Real-ISR) seeks to recover HR images from LR inputs with mixed, unknown degradations. While diffusion models surpass GANs in perceptual quality, they under-reconstruct high-frequency (HF) det…

Image Super-Resolution

Physics-Informed Super-Resolution Diffusion for 6D Phase Space Diagnostics

2025-01-08 · Alexander Scheinker

Adaptive physics-informed super-resolution diffusion is developed for non-invasive virtual diagnostics of the 6D phase space density of charged particle beams. An adaptive variational autoencoder (VAE) embeds initial bea…

Super-Resolution

Multimodal Cross-Document Event Coreference Resolution Using Linear Semantic Transfer and Mixed-Modality Ensembles

2024-04-13 · Abhijnan Nath, Huma Jamil, Shafiuddin Rehan Ahmed, George Baker 외

Event coreference resolution (ECR) is the task of determining whether distinct mentions of events within a multi-document corpus are actually linked to the same underlying occurrence. Images of the events can help facili…

coreference-resolutionCoreference ResolutionEvent Coreference Resolution