paper-with-me

Papers

IRGPT: Understanding Real-world Infrared Image with Bi-cross-modal Curriculum on Large-scale Benchmark

2025-07-19 · Zhe Cao, Jin Zhang, Ruiheng Zhang arxiv

Real-world infrared imagery presents unique challenges for vision-language models due to the scarcity of aligned text data and domain-specific characteristics. Although existing methods have advanced the field, their reliance on synthetic infrared images generated through style transfer from visible images, which limits their ability to capture the unique characteristics of the infrared modality. To address this, we propose IRGPT, the first multi-modal large language model for real-world infrared images, built upon a large-scale InfraRed-Text Dataset (IR-TD) comprising over 260K authentic image-text pairs. The proposed IR-TD dataset contains real infrared images paired with meticulously handcrafted texts, where the initial drafts originated from two complementary processes: (1) LLM-generated descriptions of visible images, and (2) rule-based descriptions of annotations. Furthermore, we introduce a bi-cross-modal curriculum transfer learning strategy that systematically transfers knowledge from visible to infrared domains by considering the difficulty scores of both infrared-visible and infrared-text. Evaluated on a benchmark of 9 tasks (e.g., recognition, grounding), IRGPT achieves state-of-the-art performance even compared with larger-scale models.

📄 PDF Abstract BibTeX arXiv:2507.14449

Code (0)

등록된 구현이 없습니다.

Tasks

Transfer LearningStyle Transfer

Similar Papers 제목 키워드 기반

HairGPT: Strand-as-Language Autoregressive Modeling for Realistic 3D Hairstyle Synthesis

2026-05-09 · Haimin Luo, Min Ouyang, Lan Xu, Jingyi Yu arxiv

Hair is a rich medium of visual and cultural expression, yet its digital modeling remains challenging due to the duality of fluidity and structure. Many existing generative approaches rely primarily on continuous diffusi…

Fair Summarization: Bridging Quality and Diversity in Extractive Summaries

2024-11-12 · Sina Bagheri Nezhad, Sayan Bandyapadhyay, Ameeta Agrawal

Fairness in multi-document summarization of user-generated content remains a critical challenge in natural language processing (NLP). Existing summarization methods often fail to ensure equitable representation across di…

DiversityDocument SummarizationExtractive SummarizationFairness+1

Toward Real-world Infrared Image Super-Resolution: A Unified Autoregressive Framework and Benchmark Dataset

2026-03-05 · Yang Zou, Jun Ma, Zhidong Jiao, Xingyuan Li 외 arxiv

Infrared image super-resolution (IISR) under real-world conditions is a practically significant yet rarely addressed task. Pioneering works are often trained and evaluated on simulated datasets or neglect the intrinsic d…

Infrared image super-resolution

Revealing Physical-World Semantic Vulnerabilities: Universal Adversarial Patches for Infrared Vision-Language Models

2026-04-03 · Chengyin Hu, Yuxian Dong, Yikun Guo, Xiang Chen 외 arxiv

Infrared vision-language models (IR-VLMs) have emerged as a promising paradigm for multimodal perception in low-visibility environments, yet their robustness to adversarial attacks remains largely unexplored. Existing ad…

MonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP and VLM Adaptation

2026-07-07 · Jiaju Han, Ma Yaqi, Yahui Chai, Xuemeng Sun 외 arxiv

Infrared remote-sensing imagery captures intensity structure, object-background contrast, and illumination-invariant cues often invisible in RGB imagery. Yet, most remote-sensing vision-language resources and models focu…