paper-with-me

홈 › Papers

Breaking Alignment Barriers: TPS-Driven Semantic Correlation Learning for Alignment-Free RGB-T Salient Object Detection

2025-12-26 · Lupiao Hu, Fasheng Wang, Fangmei Chen, Fuming Sun, Haojie Li arxiv

Existing RGB-T salient object detection methods predominantly rely on manually aligned and annotated datasets, struggling to handle real-world scenarios with raw, unaligned RGB-T image pairs. In practical applications, due to significant cross-modal disparities such as spatial misalignment, scale variations, and viewpoint shifts, the performance of current methods drastically deteriorates on unaligned datasets. To address this issue, we propose an efficient RGB-T SOD method for real-world unaligned image pairs, termed Thin-Plate Spline-driven Semantic Correlation Learning Network (TPS-SCL). We employ a dual-stream MobileViT as the encoder, combined with efficient Mamba scanning mechanisms, to effectively model correlations between the two modalities while maintaining low parameter counts and computational overhead. To suppress interference from redundant background information during alignment, we design a Semantic Correlation Constraint Module (SCCM) to hierarchically constrain salient features. Furthermore, we introduce a Thin-Plate Spline Alignment Module (TPSAM) to mitigate spatial discrepancies between modalities. Additionally, a Cross-Modal Correlation Module (CMCM) is incorporated to fully explore and integrate inter-modal dependencies, enhancing detection performance. Extensive experiments on various datasets demonstrate that TPS-SCL attains state-of-the-art (SOTA) performance among existing lightweight SOD methods and outperforms mainstream RGB-T SOD approaches.

📄 PDF Abstract BibTeX arXiv:2512.21856

Code (0)

등록된 구현이 없습니다.

Tasks

Salient Object Detection

Similar Papers 제목 키워드 기반

GOMA: Toward Structure-Driven Multimodal Alignment from a Graph Signal Smoothing Perspective

2026-05-15 · Xu Wang, Xunkai Li, Yinlin Zhu, Rong-Hua Li 외 arxiv

Multimodal alignment is commonly learned from isolated image-text pairs via CLIP-style dual encoders, leaving the relational context among entities largely unused. Multimodal attributed graphs (MAGs), where nodes carry m…

Chain-of-Lure: A Synthetic Narrative-Driven Approach to Compromise Large Language Models

2025-05-23 · Wenhan Chang, Tianqing Zhu, Yu Zhao, Shuangyong Song 외

In the era of rapid generative AI development, interactions between humans and large language models face significant misusing risks. Previous research has primarily focused on black-box scenarios using human-guided prom…

RECAP: Reproducing Copyrighted Data from LLMs Training with an Agentic Pipeline

2025-10-29 · André V. Duarte, Xuying li, Bin Zeng, Arlindo L. Oliveira 외 arxiv

If we cannot inspect the training data of a large language model (LLM), how can we ever know what it has seen? We believe the most compelling evidence arises when the model itself freely reproduces the target content. As…

Pcc-tuning: Breaking the Contrastive Learning Ceiling in Semantic Textual Similarity

2024-06-14 · BoWen Zhang, Chunping Li

Semantic Textual Similarity (STS) constitutes a critical research direction in computational linguistics and serves as a key indicator of the encoding capabilities of embedding models. Driven by advances in pre-trained l…

Contrastive LearningSemantic Textual SimilaritySentenceSTS

Breaking Barriers in Software Testing: The Power of AI-Driven Automation

2025-08-22 · Saba Naqvi, Mohammad Baqar arxiv

Software testing remains critical for ensuring reliability, yet traditional approaches are slow, costly, and prone to gaps in coverage. This paper presents an AI-driven framework that automates test case generation and v…

Reinforcement Learning