paper-with-me

홈 › Papers

Frame2Freq: Spectral Adapters for Fine-Grained Video Understanding

2026-02-21 · Thinesh Thiyakesan Ponbagavathi, Constantin Seibold, Alina Roitberg arxiv

Adapting image-pretrained backbones to video typically relies on time-domain adapters tuned to a single temporal scale. Our experiments show that these modules pick up static image cues and very fast flicker changes, while overlooking medium-speed motion. Capturing dynamics across multiple time-scales is, however, crucial for fine-grained temporal analysis (i.e., opening vs. closing bottle). To address this, we introduce Frame2Freq -- a family of frequency-aware adapters that perform spectral encoding during image-to-video adaptation of pretrained Vision Foundation Models (VFMs), improving fine-grained action recognition. Frame2Freq uses Fast Fourier Transform (FFT) along time and learns frequency-band specific embeddings that adaptively highlight the most discriminative frequency ranges. Across five fine-grained activity recognition datasets, Frame2Freq outperforms prior PEFT methods and even surpasses fully fine-tuned models on four of them. These results provide encouraging evidence that frequency analysis methods are a powerful tool for modeling temporal dynamics in image-to-video transfer. Code is available at https://github.com/th-nesh/Frame2Freq.

📄 PDF Abstract BibTeX arXiv:2602.18977

Code (0)

등록된 구현이 없습니다.

Tasks

Activity RecognitionAction Recognition

Similar Papers 제목 키워드 기반

F-Adapter: Frequency-Adaptive Parameter-Efficient Fine-Tuning in Scientific Machine Learning

2025-09-27 · Hangwei Zhang, Chun Kang, Yan Wang, Difan Zou arxiv

Parameter-efficient fine-tuning (PEFT) of powerful pre-trained models for complex downstream tasks has proven effective in vision and language processing, yet this paradigm remains unexplored in scientific machine learni…

parameter-efficient fine-tuning

SRENet: Spectral Re-Entry Network for Point Cloud Action Recognition

2026-06-02 · Qiuxia Wu, Jiarui Lan, Wenxiong Kang, Zhiyong Wang 외 arxiv

Recognizing human actions from point cloud sequences is critical for 3D perception driven applications such as autonomous driving and human-computer interaction. However, the irregular structure and temporal inconsistenc…

Representation LearningAction UnderstandingAction RecognitionAutonomous Driving

STS-Mixer: Spatio-Temporal-Spectral Mixer for 4D Point Cloud Video Understanding

2026-04-13 · Wenhao Li, Xueying Jiang, Gongjie Zhang, Xiaoqin Zhang 외 arxiv

4D point cloud videos capture rich spatial and temporal dynamics of scenes which possess unique values in various 4D understanding tasks. However, most existing methods work in the spatiotemporal domain where the underly…

Representation Learning3D Action RecognitionSemantic Segmentation

SpectralMamba-UNet: Frequency-Disentangled State Space Modeling for Texture-Structure Consistent Medical Image Segmentation

2026-02-26 · Fuhao Zhang, Lei Liu, Jialin Zhang, Ya-Nan Zhang 외 arxiv

Accurate medical image segmentation requires effective modeling of both global anatomical structures and fine-grained boundary details. Recent state space models (e.g., Vision Mamba) offer efficient long-range dependency…

Medical Image Segmentation

Time-Frequency Weighted Losses for Phoneme Reconstruction in DNN-Based Speech Enhancement

2026-06-19 · Nasser-Eddine Monir, Paul Magron, Romain Serizel arxiv

Conventional training losses for speech enhancement based on the signal-to-distortion ratio (SDR) treat all time-frequency (TF) regions uniformly, overlooking the fine-grained spectral cues that are relevant to specific …

Speech Enhancement