paper-with-me

홈 › Papers

FLARE: Fast Low-rank Attention Routing Engine

2025-08-18 · Vedant Puri, Aditya Joglekar, Sri Datta Ganesh Bandreddi, Kevin Ferguson, Yu-hsuan Chen, Yongjie Jessica Zhang, Levent Burak Kara arxiv

The quadratic complexity of self-attention limits the scalability of transformers on long sequences. We introduce Fast Low-rank Attention Routing Engine (FLARE), a token-mixing operator that realizes low-rank attention by routing information through a small set of latent tokens. Each layer induces an input-input token mixing matrix of rank at most $M$ via a minimal encode-decode factorization implemented using only two standard scaled dot-product attention (SDPA) calls. Because the dominant ${O}(NM)$ computation is expressed purely in terms of standard SDPA, FLARE is compatible with fused attention kernels and avoids materializing $M\times N$ projection matrices. FLARE further assigns disjoint latent slices to each attention head, yielding a mixture of head-specific low-rank pathways. Empirically, FLARE scales to one-million-point unstructured meshes on a single GPU, achieves state-of-the-art accuracy on PDE surrogate benchmarks, and outperforms general-purpose efficient-attention methods on the Long Range Arena suite. We additionally release a large-scale additive manufacturing benchmark dataset. Our code is available at https://github.com/vpuri3/FLARE.py.

📄 PDF Abstract BibTeX arXiv:2508.12594

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

High-Fidelity Industrial Crash Dynamics Prediction via Geometry-Aware Operator Learning with Memory-Efficient Low-Rank Attention

2026-05-26 · Deepak Akhare, Mohammad Amin Nabian, Corey Adams, Sudeep Chavare 외 arxiv

Automotive crashworthiness optimization remains a safety-critical challenge, requiring the management of large-scale nonlinear structural deformations and energy dissipation through iterative, high-fidelity simulations. …

Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models

2026-06-10 · Cheng-Yu Yang, Shao-Yuan Lo, Yu-Lun Liu arxiv

Vision-language models (VLMs) project images into hundreds to thousands of visual tokens, making decoder inference expensive in both attention computation and KV-cache memory. Existing visual-token reduction methods larg…

FLARE MCMC: Fidelity-based Layer-Adaptive REcursive proposals for MCMC

2026-08-13 · Harini Venkatesan, Christian Shelton, Ming-Feng Ho, Simeon Bird 외 arxiv

Markov chain Monte Carlo (MCMC) requires only the ability to evaluate the likelihood, making it a common technique for inference in complex models. However, it can have a slow mixing rate, requiring the generation of man…

FF-Former: Swin Fourier Transformer for Nighttime Flare Removal

2023-03-20 · CVPR Workshop 2023 3 · Dafeng Zhang, Jia Ouyang, Guanqun Liu, Xiaobing Wang 외

In the process of removing nighttime flare, it is crucial to have a large receptive field due to the fact that flare can occupy a substantial portion of an image, even potentially the entire image. However, the conventio…

Flare Removal

FlareX: A Physics-Informed Dataset for Lens Flare Removal via 2D Synthesis and 3D Rendering

2025-10-11 · Lishen Qu, Zhihao Liu, Jinshan Pan, Shihao Zhou 외 arxiv

Lens flare occurs when shooting towards strong light sources, significantly degrading the visual quality of images. Due to the difficulty in capturing flare-corrupted and flare-free image pairs in the real world, existin…

Flare Removal