paper-with-me

Papers

Breaking Scale Anchoring: Frequency Representation Learning for Accurate High-Resolution Inference from Low-Resolution Training

2025-11-28 · Wenshuo Wang, Fan Zhang arxiv

Zero-Shot Super-Resolution Spatiotemporal Forecasting requires a deep learning model to be trained on low-resolution data and deployed for inference on high-resolution. Existing studies consider maintaining similar error across different resolutions as indicative of successful multi-resolution generalization. However, deep learning models serving as alternatives to numerical solvers should reduce error as resolution increases. The fundamental limitation is, the upper bound of physical law frequencies that low-resolution data can represent is constrained by its Nyquist frequency, making it difficult for models to process signals containing unseen frequency components during high-resolution inference. This results in errors being anchored at low resolution, incorrectly interpreted as successful generalization. We define this fundamental phenomenon as a new problem distinct from existing issues: Scale Anchoring. Therefore, we propose architecture-agnostic Frequency Representation Learning. It alleviates Scale Anchoring through resolution-aligned frequency representations and spectral consistency training: on grids with higher Nyquist frequencies, the frequency response in high-frequency bands of FRL-enhanced variants is more stable. This allows errors to decrease with resolution and significantly outperform baselines within our task and resolution range, while incurring only modest computational overhead.

📄 PDF Abstract BibTeX arXiv:2512.05132

Code (0)

등록된 구현이 없습니다.

Tasks

Representation Learning

Similar Papers 제목 키워드 기반

STAND: Semantic Anchoring Constraint with Dual-Granularity Disambiguation for Remote Sensing Image Change Captioning

2026-04-25 · Yanpei Gong, Beichen Zhang, Hao Wang, Zhaobo Qi 외 arxiv

Remote sensing image change captioning (RSICC) aims to describe the difference between two remote sensing images. While recent methods have explored video modeling, they largely overlook the inherent ambiguities in viewp…

From Noise to Intent: Anchoring Generative VLA Policies with Residual Bridges

2026-04-23 · Yiming Zhong, Yaoyu He, Zemin Yang, Pengfei Tian 외 arxiv

Bridging high-level semantic understanding with low-level physical control remains a persistent challenge in embodied intelligence, stemming from the fundamental spatiotemporal scale mismatch between cognition and action…

RCL: Recurrent Continuous Localization for Temporal Action Detection

2022-03-14 · CVPR 2022 1 · Qiang Wang, Yanhao Zhang, Yun Zheng, Pan Pan

Temporal representation is the cornerstone of modern action detection techniques. State-of-the-art methods mostly rely on a dense anchoring scheme, where anchors are sampled uniformly over the temporal domain with a disc…

Action Detection

FORCE: Transferable Visual Jailbreaking Attacks via Feature Over-Reliance CorrEction

2025-09-25 · Runqi Lin, Alasdair Paren, Suqin Yuan, Muyang Li 외 arxiv

The integration of new modalities enhances the capabilities of multimodal large language models (MLLMs) but also introduces additional vulnerabilities. In particular, simple visual jailbreaking attacks can manipulate ope…

What is in Your Safe Data? Identifying Benign Data that Breaks Safety

2024-04-01 · Luxi He, Mengzhou Xia, Peter Henderson

Current Large Language Models (LLMs), even those tuned for safety and alignment, are susceptible to jailbreaking. Some have found that just further fine-tuning an aligned model with benign data (i.e., data without harmfu…

Math