paper-with-me

홈 › Papers

Can a Second-View Image Be a Language? Geometric and Semantic Cross-Modal Reasoning for X-ray Prohibited Item Detection

2025-11-23 · Chuang Peng, Renshuai Tao, Zhongwei Ren, Xianglong Liu, Yunchao Wei arxiv

Automatic X-ray prohibited items detection is vital for security inspection and has been widely studied. Traditional methods rely on visual modality, often struggling with complex threats. While recent studies incorporate language to guide single-view images, human inspectors typically use dual-view images in practice. This raises the question: can the second view provide constraints similar to a language modality? In this work, we introduce DualXrayBench, the first comprehensive benchmark for X-ray inspection that includes multiple views and modalities. It supports eight tasks designed to test cross-view reasoning. In DualXrayBench, we introduce a caption corpus consisting of 45,613 dual-view image pairs across 12 categories with corresponding captions. Building upon these data, we propose the Geometric (cross-view)-Semantic (cross-modality) Reasoner (GSR), a multimodal model that jointly learns correspondences between cross-view geometry and cross-modal semantics, treating the second-view images as a "language-like modality". To enable this, we construct the GSXray dataset, with structured Chain-of-Thought sequences: <top>, <side>, <conclusion>. Comprehensive evaluations on DualXrayBench demonstrate that GSR achieves significant improvements across all X-ray tasks, offering a new perspective for real-world X-ray inspection.

📄 PDF Abstract BibTeX arXiv:2511.18385

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

EAGLE: Efficient Adaptive Geometry-based Learning in Cross-view Understanding

2024-06-03 · Thanh-Dat Truong, Utsav Prabhu, Dongyi Wang, Bhiksha Raj 외

Unsupervised Domain Adaptation has been an efficient approach to transferring the semantic segmentation model across data distributions. Meanwhile, the recent Open-vocabulary Semantic Scene understanding based on large-s…

Domain AdaptationOpen Vocabulary Semantic SegmentationOpen-Vocabulary Semantic SegmentationScene Understanding+3

GitNet: Geometric Prior-based Transformation for Birds-Eye-View Segmentation

2022-04-16 · Shi Gong, Xiaoqing Ye, Xiao Tan, Jingdong Wang 외

Birds-eye-view (BEV) semantic segmentation is critical for autonomous driving for its powerful spatial representation ability. It is challenging to estimate the BEV semantic maps from monocular images due to the spatial …

Autonomous DrivingBEV SegmentationImage SegmentationSegmentation+1

OV-NeRF: Open-vocabulary Neural Radiance Fields with Vision and Language Foundation Models for 3D Semantic Understanding

2024-02-07 · Guibiao Liao, Kaichen Zhou, Zhenyu Bao, Kanglin Liu 외

The development of Neural Radiance Fields (NeRFs) has provided a potent representation for encapsulating the geometric and appearance characteristics of 3D scenes. Enhancing the capabilities of NeRFs in open-vocabulary 3…

NeRF

Cog3DMap: Multi-View Vision-Language Reasoning with 3D Cognitive Maps

2026-03-24 · Chanyoung Gwak, Yoonwoo Jeong, Byungwoo Jeon, Hyunseok Lee 외 arxiv

Precise spatial understanding from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs), as their visual representations are predominantly semantic and lack explicit geometric gr…

Spatial Reasoning

SECOND-Grasp: Semantic Contact-guided Dexterous Grasping

2026-05-13 · Han Yi Shin, Heeju Ko, Jaewon Mun, Qixing Huang 외 arxiv

Achieving reliable robotic manipulation, such as dexterous grasping, requires a synergy between physically stable interactions and semantic task guidance, yet these objectives are often treated as separate, disjoint goal…