paper-with-me

홈 › Papers

SPIRIT: Adapting Vision Foundation Models for Unified Single- and Multi-Frame Infrared Small Target Detection

2026-02-02 · Qian Xu, Xi Li, Fei Gao, Jie Guo, Haojuan Yuan, Shuaipeng Fan, Mingjin Zhang arxiv

Infrared small target detection (IRSTD) is crucial for surveillance and early-warning, with deployments spanning both single-frame analysis and video-mode tracking. A practical solution should leverage vision foundation models (VFMs) to mitigate infrared data scarcity, while adopting a memory-attention-based temporal propagation framework that unifies single- and multi-frame inference. However, infrared small targets exhibit weak radiometric signals and limited semantic cues, which differ markedly from visible-spectrum imagery. This modality gap makes direct use of semantics-oriented VFMs and appearance-driven cross-frame association unreliable for IRSTD: hierarchical feature aggregation can submerge localized target peaks, and appearance-only memory attention becomes ambiguous, leading to spurious clutter associations. To address these challenges, we propose SPIRIT, a unified and VFM-compatible framework that adapts VFMs to IRSTD via lightweight physics-informed plug-ins. Spatially, PIFR refines features by approximating rank-sparsity decomposition to suppress structured background components and enhance sparse target-like signals. Temporally, PGMA injects history-derived soft spatial priors into memory cross-attention to constrain cross-frame association, enabling robust video detection while naturally reverting to single-frame inference when temporal context is absent. Experiments on multiple IRSTD benchmarks show consistent gains over VFM-based baselines and SOTA performance.

📄 PDF Abstract BibTeX arXiv:2602.01843

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SAM3D: Zero-Shot 3D Object Detection via Segment Anything Model

2023-06-04 · Dingyuan Zhang, Dingkang Liang, Hongcheng Yang, Zhikang Zou 외

With the development of large language models, many remarkable linguistic systems like ChatGPT have thrived and achieved astonishing success on many tasks, showing the incredible power of foundation models. In the spirit…

3D Object DetectionImage SegmentationObjectobject-detection+2

MMSense: Adapting Vision-based Foundation Model for Multi-task Multi-modal Wireless Sensing

2025-11-15 · Zhizhen Li, Xuanhao Luo, Xueren Ge, Longyu Zhou 외 arxiv

Large AI models have been widely adopted in wireless communications for channel modeling, beamforming, and resource optimization. However, most existing efforts remain limited to single-modality inputs and channel-specif…

Foundation Models for Face Presentation Attack Detection: A Unified Linear-Probing Benchmark

2026-07-29 · Peter Lorenz, Anjith George, Sébastien Marcel arxiv

Face presentation attack detection (PAD) remains challenging under cross-dataset evaluation, where domain shift degrades models trained on a single dataset. The scarcity of large-scale labeled data motivates adapting pre…

Face Presentation Attack Detection

Spirit LM: Interleaved Spoken and Written Language Model

2024-02-08 · Tu Anh Nguyen, Benjamin Muller, Bokai Yu, Marta R. Costa-Jussa 외

We introduce Spirit LM, a foundation multimodal language model that freely mixes text and speech. Our model is based on a 7B pretrained text language model that we extend to the speech modality by continuously training i…

Language ModelingLanguage Modellingmodel

ABACUS: Adapting Unified Foundation Model for Bridging Image Count Understanding and Generation

2026-06-22 · Anindya Mondal, Sauradip Nag, Anjan Dutta arxiv

ABACUS is a unified vision-language model that handles object counting, crowd counting, referring-expression counting, and count-faithful image generation without any benchmark-specific training required. Our model is bu…

Object LocalizationImage GenerationObject CountingCrowd Counting