paper-with-me

홈 › Papers

Beyond Similarity: Temporal Operator Attention for Time Series Analysis

2026-05-11 · Jevon Twitty, Vinh Pham, Nitiwith Rotchanarak, Viresh Pati, Yubin Kim, Shihao Yang, Jiecheng Lu arxiv

A persistent paradox in time-series forecasting is that structurally simple MLP and linear models often outperform high-capacity Transformers. We argue that this gap arises from a mismatch in the sequence-modeling primitive: while many time-series dynamics are governed by global temporal operators (e.g., filtering and harmonic structure), standard attention forms each output as a convex combination of inputs. This restricts its ability to represent signed and oscillatory transformations that are fundamental to temporal signal processing. We formalize this limitation as a simplex-constrained mixing bottleneck in softmax attention, which becomes especially restrictive for operator-driven time-series tasks. To address this, we propose $\textbf{Temporal Operator Attention (TOA)}$, a framework that augments attention with explicit, learnable sequence-space operators, enabling direct signed mixing across time while preserving input-dependent adaptivity. To make dense $N \times N$ operators practical, we introduce Stochastic Operator Regularization, a high-variance dropout mechanism that stabilizes training and prevents trivial memorization. Across forecasting, anomaly detection, and classification benchmarks, TOA consistently improves performance when integrated into standard backbones such as PatchTST and iTransformer, with particularly strong gains in reconstruction-heavy tasks. These results suggest that explicit operator learning is a key ingredient for effective time-series modeling.

📄 PDF Abstract BibTeX arXiv:2605.11287

Code (0)

등록된 구현이 없습니다.

Tasks

Time Series AnalysisAnomaly Detection

Similar Papers 제목 키워드 기반

Multiscale Self Attentive Convolutions for Vision and Language Modeling

2019-12-03 · Oren Barkan

Self attention mechanisms have become a key building block in many state-of-the-art language understanding models. In this paper, we show that the self attention operator can be formulated in terms of 1x1 convolution ope…

Language ModelingLanguage Modelling

Efficient Temporal Modeling for Mobile Sleep Staging via Lightweight Random Attention

2026-05-31 · Guisong Liu, Pengfei Wei, Jainsong Zhang, Martin Dresler arxiv

Mobile sleep staging serves as a foundational infrastructure for in-home sleep monitoring and closed-loop modulation. But existing sequential models such as RNNs and Transformers are computationally expensive for mobile …

Siamese Attention Networks

2019-09-25 · Hongyang Gao, Yaochen Xie, Shuiwang Ji

Attention operators have been widely applied on data of various orders and dimensions such as texts, images, and videos. One challenge of applying attention operators is the excessive usage of computational resources. Th…

image-classificationImage Classification

DiTFastAttn: Attention Compression for Diffusion Transformer Models

2024-06-12 · Zhihang Yuan, Hanling Zhang, Pu Lu, Xuefei Ning 외

Diffusion Transformers (DiT) excel at image and video generation but face computational challenges due to the quadratic complexity of self-attention operators. We propose DiTFastAttn, a post-training compression method t…

2kImage GenerationVideo Generation

Learning Physical Operators using Neural Operators

2026-02-26 · Vignesh Gopakumar, Ander Gray, Dan Giles, Lorenzo Zanisi 외 arxiv

Neural operators have emerged as promising surrogate models for solving partial differential equations (PDEs), but struggle to generalise beyond training distributions and are often constrained to a fixed temporal discre…