paper-with-me

홈 › Papers

A Comparison of Object Detection and Phrase Grounding Models in Chest X-ray Abnormality Localization using Eye-tracking Data

2025-03-02 · Elham Ghelichkhan, Tolga Tasdizen

Chest diseases rank among the most prevalent and dangerous global health issues. Object detection and phrase grounding deep learning models interpret complex radiology data to assist healthcare professionals in diagnosis. Object detection locates abnormalities for classes, while phrase grounding locates abnormalities for textual descriptions. This paper investigates how text enhances abnormality localization in chest X-rays by comparing the performance and explainability of these two tasks. To establish an explainability baseline, we proposed an automatic pipeline to generate image regions for report sentences using radiologists' eye-tracking data. The better performance - mIoU = 0.36 vs. 0.20 - and explainability - Containment ratio 0.48 vs. 0.26 - of the phrase grounding model infers the effectiveness of text in enhancing chest X-ray abnormality localization.

📄 PDF Abstract BibTeX arXiv:2503.01037

Code (0)

등록된 구현이 없습니다.

Tasks

object-detectionObject DetectionPhrase Grounding

Similar Papers 제목 키워드 기반

LoFi: Location-Aware Fine-Grained Representation Learning for Chest X-ray

2026-03-19 · Myeongkyun Kang, Yanting Yang, Xiaoxiao Li arxiv

Fine-grained representation learning is crucial for retrieval and phrase grounding in chest X-rays, where clinically relevant findings are often spatially confined. However, the lack of region-level supervision in contra…

Representation LearningDense CaptioningPhrase Grounding

TRACE: Temporal Radiology with Anatomical Change Explanation for Grounded X-ray Report Generation

2026-02-03 · OFM Riaz Rahman Aranya, Kevin Desai arxiv

Temporal comparison of chest X-rays is fundamental to clinical radiology, enabling detection of disease progression, treatment response, and new findings. While vision-language models have advanced single-image report ge…

Change DetectionVisual Grounding

DQ-DETR: Dual Query Detection Transformer for Phrase Extraction and Grounding

2022-11-28 · Shilong Liu, Yaoyuan Liang, Feng Li, Shijia Huang 외

In this paper, we study the problem of visual grounding by considering both phrase extraction and grounding (PEG). In contrast to the previous phrase-known-at-test setting, PEG requires a model to extract phrases from te…

object-detectionObject DetectionPhrase Extraction and Grounding (PEG)Phrase Grounding+3

Utilizing Every Image Object for Semi-supervised Phrase Grounding

2020-11-05 · Haidong Zhu, Arka Sadhu, Zhaoheng Zheng, Ram Nevatia

Phrase grounding models localize an object in the image given a referring expression. The annotated language queries available during training are limited, which also limits the variations of language combinations that a…

Phrase GroundingReferring Expression

AnatomiX, an Anatomy-Aware Grounded Multimodal Large Language Model for Chest X-Ray Interpretation

2026-01-06 · Anees Ur Rehman Hashmi, Numan Saeed, Christoph Lippert arxiv

Multimodal medical large language models have shown substantial progress in chest X-ray interpretation but continue to face challenges in spatial reasoning and anatomical understanding. Although existing grounding techni…

Visual Question AnsweringSpatial ReasoningPhrase Grounding