paper-with-me

홈 › Papers

LoCA: Spatially-Aware Low-Rank Convolutional Adaptation of Vision Foundation Models

2026-07-08 · Sojung An, Junha Lee, Sujeong You, Nam Ik Cho, Donghyun Kim arxiv

Pre-trained Vision Foundation Models (VFMs) provide strong visual representations for diverse downstream tasks. The key challenge of VFM adaptation stems from the prohibitive costs of full fine-tuning and catastrophic forgetting. To address this, Low-Rank Adaptation (LoRA) has emerged as the prevailing paradigm for Parameter-Efficient Fine-Tuning (PEFT). However, LoRA is typically designed for transformer self-attention layers parameterized by 2D matrices. Since convolutional kernels inherently couple spatial and channel information within a 4D tensor, forcing them into a monolithic 2D matrix disrupts the inherent spatial topology. In this paper, we propose Low-Rank Convolutional Adaptation (LoCA), a convolution-aware PEFT framework that addresses spatial-channel entanglement by decoupling channel and spatial adaptation. LoCA introduces a low-rank channel adaptation for dense cross-channel mixing and refines spatial bases extracted from pre-trained kernels via Singular Value Decomposition (SVD). Experimental results show that LoCA preserves pre-trained spatial priors and achieves competitive or state-of-the-art performance across fine-grained classification, domain-generalized semantic segmentation, and generative benchmarks.

📄 PDF Abstract BibTeX arXiv:2607.06918

Code (3)

NickDee96/ASR-TTS-paper-daily ★ 3
Tavish9/awesome-daily-AI-arxiv ★ 111
ZhuYingJessica/cv-daily ★ 47

Tasks

parameter-efficient fine-tuningSemantic Segmentation

Similar Papers 제목 키워드 기반

Rethinking Autoregressive Models for Lossless Image Compression via Hierarchical Parallelism and Progressive Adaptation

2025-11-14 · Daxin Li, Yuanchao Bai, Kai Wang, Wenbo Zhao 외 arxiv

Autoregressive (AR) models, the theoretical performance benchmark for learned lossless image compression, are often dismissed as impractical due to prohibitive computational cost. This work re-thinks this paradigm, intro…

Image Compression

GLoRIA: Gated Low-Rank Interpretable Adaptation for Dialectal ASR

2026-03-02 · Pouya Mehralian, Melissa Farasyn, Anne Breitbarth, Anne-Sophie Ghyselen 외 arxiv

Automatic Speech Recognition (ASR) in dialect-heavy settings remains challenging due to strong regional variation and limited labeled data. We propose GLoRIA, a parameter-efficient adaptation framework that leverages met…

Speech Recognition

LoCA: Location-Aware Cosine Adaptation for Parameter-Efficient Fine-Tuning

2025-02-05 · Zhekai Du, Yinjie Min, Jingjing Li, Ke Lu 외

Low-rank adaptation (LoRA) has become a prevalent method for adapting pre-trained large language models to downstream tasks. However, the simple low-rank decomposition form may constrain the hypothesis space. To address …

parameter-efficient fine-tuning

CoordGate: Efficiently Computing Spatially-Varying Convolutions in Convolutional Neural Networks

2024-01-09 · Sunny Howard, Peter Norreys, Andreas Döpp

Optical imaging systems are inherently limited in their resolution due to the point spread function (PSF), which applies a static, yet spatially-varying, convolution to the image. This degradation can be addressed via Co…

DeblurringImage Deblurring

Spatially Aware Linear Transformer (SAL-T) for Particle Jet Tagging

2025-10-24 · Aaron Wang, Zihan Zhao, Subash Katel, Vivekanand Gyanchand Sahu 외 arxiv

Transformers are very effective in capturing both global and local correlations within high-energy particle collisions, but they present deployment challenges in high-data-throughput environments, such as the CERN LHC. T…

Point Cloud ClassificationJet Tagging