paper-with-me

홈 › Papers

Structured Spectral Reasoning for Frequency-Adaptive Multimodal Recommendation

2025-12-01 · Wei Yang, Rui Zhong, Yiqun Chen, Chi Lu, Peng Jiang arxiv

Multimodal recommendation aims to integrate collaborative signals with heterogeneous content such as visual and textual information, but remains challenged by modality-specific noise, semantic inconsistency, and unstable propagation over user-item graphs. These issues are often exacerbated by naive fusion or shallow modeling strategies, leading to degraded generalization and poor robustness. While recent work has explored the frequency domain as a lens to separate stable from noisy signals, most methods rely on static filtering or reweighting, lacking the ability to reason over spectral structure or adapt to modality-specific reliability. To address these challenges, we propose a Structured Spectral Reasoning (SSR) framework for frequency-aware multimodal recommendation. Our method follows a four-stage pipeline: (i) Decompose graph-based multimodal signals into spectral bands via graph-guided transformations to isolate semantic granularity; (ii) Modulate band-level reliability with spectral band masking, a training-time masking with a prediction-consistency objective that suppresses brittle frequency components; (iii) Fuse complementary frequency cues using hyperspectral reasoning with low-rank cross-band interaction; and (iv) Align modality-specific spectral features via contrastive regularization to promote semantic and structural consistency. Experiments on three real-world benchmarks show consistent gains over strong baselines, particularly under sparse and cold-start settings. Additional analyses indicate that structured spectral modeling improves robustness and provides clearer diagnostics of how different bands contribute to performance.

📄 PDF Abstract BibTeX arXiv:2512.01372

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal Recommendation

Similar Papers 제목 키워드 기반

A Spatial-Spectral-Frequency Interactive Network for Multimodal Remote Sensing Classification

2025-10-06 · Hao Liu, Yunhao Gao, Wei Li, Mingyang Zhang 외 arxiv

Deep learning-based methods have achieved significant success in remote sensing Earth observation data analysis. Numerous feature fusion techniques address multimodal remote sensing image classification by integrating gl…

Remote Sensing Image Classification

Delving into Non-Exchangeability for Conformal Prediction in Graph-Structured Multivariate Time Series

2026-05-06 · Ruichao Guo, Xingyao Han, Luo Wenshui, Zhe Liu 외 arxiv

Point forecasting for graph-structured multivariate time series is a fundamental problem, but rigorous uncertainty quantification for such predictions is still underexplored. Conformal prediction (CP) offers uncertainty …

CHASM: Cross-frequency Harmonized Axis-Separable Mixing for Spectral Token Operators

2026-05-14 · Pengcheng Fang, Hongli Chen, Yuxia Chen, Tengjiao Sun 외 arxiv

Spectral token mixers based on Fourier transforms provide an efficient way to model global interactions in visual feature maps. Existing designs often either apply filter-wise spectral responses along fixed channel axes,…

Image ReconstructionMRI Reconstruction

Spectral Evolution-Guided Token Pruning in Multimodal Large Language Models

2026-06-23 · Bin Chen, Yuxiang Cai, Yadan Luo, Yi Zhang 외 arxiv

Reducing visual token redundancy is critical for accelerating Multimodal Large Language Models (MLLMs) without degrading cross-modal reasoning performance. Existing token pruning methods typically rely on single-layer si…

Spectral-Progressive Thought Flow for Lightweight Multimodal Reasoning

2026-06-01 · Yixian Shen, Zhiheng Yang, Qi Bi, Changshuo Wang 외 arxiv

Multimodal spatial reasoning often relies on long chains of intermediate textual and visual thoughts, where accumulating visual tokens and dense cross-modal attention incur substantial computation and memory overhead. To…

Multimodal ReasoningSpatial Reasoning