paper-with-me

홈 › Papers

Pseudo-Text-Conditioned 3D Grounding DINO for Organ Localization in Abdominal CT

2026-06-25 · Siqi Chen, Han Gong, Keyi Hou, Jingxuan Yang, Sheethal Bhat, Andreas Maier arxiv

Reliable organ localization in abdominal CT can provide spatial priors for downstream trauma analysis. We propose CT-3GDINO, a lightweight 3D detector that adapts a Grounding-DINO-style query-based architecture to fixed organ localization using frozen pseudo-text class tokens instead of a real text encoder. The model combines a Swin3D visual backbone, bidirectional feature enhancement, pseudo-text-guided query selection, and a cross-modality decoder to predict normalized 3D boxes for liver, spleen, left kidney, right kidney, and bowel. We train and evaluate on 193 matched RSNA/RATIC CT volumes with segmentation-derived boxes. The best multi-scale model, trained from scratch, achieves 0.5830 overall top-1 class-wise mAP over 3D IoU thresholds from 0.1 to 0.7, outperforming fixed- and trainable-backbone classification-pretrained variants with 0.5570 and 0.4657 mAP. Performance is strong for coarse localization, with 0.9649 AP at IoU 0.1, but remains limited for strict box alignment, with 0.1552 AP at IoU 0.7. These results establish CT-3GDINO as an open-source baseline for pseudo-text-conditioned 3D organ localization and motivate future work on localization-aware pretraining, richer multimodal conditioning, and injury-focused detection.

📄 PDF Abstract BibTeX arXiv:2606.27084

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

GroundingDINO-US-SAM: Text-Prompted Multi-Organ Segmentation in Ultrasound with LoRA-Tuned Vision-Language Models

2025-06-30 · Hamza Rasaee, Taha Koleilat, Hassan Rivaz

Accurate and generalizable object segmentation in ultrasound imaging remains a significant challenge due to anatomical variability, diverse imaging protocols, and limited annotated data. In this study, we propose a promp…

Organ SegmentationSegmentationSemantic Segmentation

Efficient License Plate Recognition via Pseudo-Labeled Supervision with Grounding DINO and YOLOv8

2025-10-28 · Zahra Ebrahimi Vargoorani, Amir Mohammad Ghoreyshi, Ching Yee Suen arxiv

Developing a highly accurate automatic license plate recognition system (ALPR) is challenging due to environmental factors such as lighting, rain, and dust. Additional difficulties include high vehicle speeds, varying ca…

License Plate RecognitionLicense Plate Detection

LoGSAM: Parameter-Efficient Cross-Modal Grounding for MRI Segmentation

2026-03-18 · Mohammad Robaitul Islam Bhuiyan, Sheethal Bhat, Melika Qahqaie arxiv

Precise localization and delineation of brain tumors using magnetic resonance imaging (MRI) are essential for planning therapy and guiding surgical decisions. To address this, we propose LoGSAM, a parameter-efficient, de…

Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection

2024-05-16 · Tianhe Ren, Qing Jiang, Shilong Liu, Zhaoyang Zeng 외

This paper introduces Grounding DINO 1.5, a suite of advanced open-set object detection models developed by IDEA Research, which aims to advance the "Edge" of open-set object detection. The suite encompasses two models: …

Edge-computingFew-Shot Object Detectionobject-detectionObject Detection+1

ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations

2025-01-24 · Tianming Liang, Kun-Yu Lin, Chaolei Tan, JianGuo Zhang 외

Referring video object segmentation (RVOS) aims to segment target objects throughout a video based on a text description. Despite notable progress in recent years, current RVOS models remain struggle to handle complicate…

DecoderObjectReferring Expression SegmentationReferring Video Object Segmentation+4