paper-with-me

Phrase Grounding

5개 벤치마크 · 논문 98편 · 이 태스크의 논문 보기 →

Benchmarks

Flickr30k

결과 3개

ReferIt

결과 3개

Visual Genome

결과 3개

Most implemented

Towards Visual Grounding: A Survey

2024-12-28 · 구현 4개

Papers

Vision-Language Grounding as Bidirectional Concept Correspondence

2026-08-08 · Jieyu Zhang, Ziqi Gao, Luke Zettlemoyer, Ranjay Krishna hf

Vision-language grounding connects language to visual content, yet most existing formulations reduce grounding to a unidirectional localization problem: given a prespecified text phrase or category name, identify the cor…

Referring ExpressionImage SegmentationPhrase Grounding

Location-Aware Fine-Grained Representation Learning for Medical Vision Foundation Models

2026-08-02 · Myeongkyun Kang, Yanting Yang, Xiaoxiao Li arxiv

Fine-grained visual representations are essential for medical image analysis, particularly when diagnostically relevant evidence is subtle and spatially localized. Modern transformer-based medical vision encoders must th…

Visual Question AnsweringSelf-Supervised LearningRepresentation LearningPhrase Grounding

LoFi: Location-Aware Fine-Grained Representation Learning for Chest X-ray

2026-03-19 · Myeongkyun Kang, Yanting Yang, Xiaoxiao Li arxiv

Fine-grained representation learning is crucial for retrieval and phrase grounding in chest X-rays, where clinically relevant findings are often spatially confined. However, the lack of region-level supervision in contra…

Representation LearningDense CaptioningPhrase Grounding

LLM-Guided Diagnostic Evidence Alignment for Medical Vision-Language Pretraining under Limited Pairing

2026-02-07 · Huimin Yan, Liang Bai, Xian Yang, Long Chen arxiv

Most existing CLIP-style medical vision--language pretraining methods rely on global or local alignment with substantial paired data. However, global alignment is easily dominated by non-diagnostic information, while loc…

Phrase GroundingText Retrieval

CURE: Curriculum-guided Multi-task Training for Reliable Anatomy Grounded Report Generation

2026-01-21 · Pablo Messina, Andrés Villa, Juan León Alcázar, Karen Sánchez 외 arxiv

Medical vision-language models can automate the generation of radiology reports but struggle with accurate visual grounding and factual consistency. Existing models often misalign textual findings with visual evidence, l…

Visual GroundingPhrase Grounding

AnatomiX, an Anatomy-Aware Grounded Multimodal Large Language Model for Chest X-Ray Interpretation

2026-01-06 · Anees Ur Rehman Hashmi, Numan Saeed, Christoph Lippert arxiv

Multimodal medical large language models have shown substantial progress in chest X-ray interpretation but continue to face challenges in spatial reasoning and anatomical understanding. Although existing grounding techni…

Visual Question AnsweringSpatial ReasoningPhrase Grounding

전체 98편 보기 →